35 hpc engineer jobs at 18 companies in Des Moines, WA

2w
Save
Mark Applied
Hide
Senior HPC Cluster Engineer
Bothell or Boulder or College Park
$146k-$209k/yr HybridFull Time
IonQ
IonQNYSE: IONQ: Develops and sells trapped-ion quantum computers and cloud services.
5+ YOEBachelor's or equivalent,5+ years managing Linux HPC clusters,experience with Slurm/PBS/Grid Engine,Git,Python/Go/Bash,Ansible,and supporting scientific workloads.
Linux, Slurm, PBS, Grid Engine, Git, Python, Go, Bash, Ansible, NVIDIA, InfiniBand, RoCE, Docker, Podman, Singularity, Enroot, Kubernetes, GCP, AWS
2mo
Save
Mark Applied
Hide
Senior HPC Cluster Engineer
Santa Clara or Austin or Redmond
$152k-$288k/yr OnsiteFull Time
NVIDIA
NVIDIANASDAQ: NVDA: Designs graphics processing units and artificial intelligence hardware.
5+ YOE5+ years operating large-scale compute infrastructure; Bachelor's degree or equivalent; experience with BCM/Ansible, Slurm/LSF/PBS/K8s, Linux, containers, Python/Bash, MPI/NCCL; EDA and performance tuning experience preferred.
BCM, Ansible, Slurm, LSF, PBS, K8s, MPI, NCCL, Rocky, Centos, RHEL, Ubuntu, Enroot, Docker, Python, Bash, CUDA, MLPerf, InfiniBand, RDMA, RoCE, Lustre, GPFS, Prometheus, OpenSearch, Grafana, NVIDIA GPUs
12h
Save
Mark Applied
Hide
HPC Operations Engineer
Bellevue or Seattle
$200k/yr OnsiteFull Time
Evergreen Statistical Trading
Evergreen Statistical Trading: Proprietary trading firm specializing in statistical research and algorithmic trading.
Strong Linux systems administration and Python automation experience, plus experience with Git, monitoring, schedulers, cluster provisioning, filesystems, and InfiniBand; able to troubleshoot critical systems.
Linux, Red Hat Enterprise Linux (RHEL), Rocky Linux, AlmaLinux, Python, Git, Prometheus, Grafana, Slurm, Warewulf, xCAT, ZFS, GPFS, Lustre, InfiniBand
1w
Save
Mark Applied
Hide
Director, HPC Systems Software Engineering
Houston or New York City or San Francisco or Seattle
$230k-$343k/yr OnsiteFull Time
Nscale
Nscale: Vertically integrated AI infrastructure provider for high-performance computing.
Requires leadership of HPC or infrastructure engineering organizations, Linux, distributed systems, production operations, GPU infrastructure, networking, budgeting, and people management; Slurm experience strongly preferred.
Slurm, Kueue, Linux
1w
Save
Mark Applied
Hide
Senior HPC Support Engineer, InfiniBand - NVLink
Westford or Durham or Redmond or Santa Clara or New York City
$108k-$207k/yr OnsiteFull Time
NVIDIA
NVIDIANASDAQ: NVDA: Designs GPU-accelerated computing and artificial intelligence hardware.
5+ YOE5+ years of customer support and debugging in large-scale networking or AI infrastructure; networking, Linux, cloud, containers, virtualization, InfiniBand, and GPU expertise; English proficiency and relevant degree or equivalent experience.
Linux, InfiniBand, NVLink, Ethernet, GPU, Claude, Codex, Cursor, IP, TCPDUMP, Wireshark, Open platforms, KVM, ESXi, AWS, OCI, RDMA/RoCEv2, NCCL, MPI, Slurm/SchedMD, Bash, Python, CCIE, JNCIE-DC/ENT, RHCE, NCP-AII/AIO/AIN
1mo
Save
Mark Applied
Hide
Site Reliability Engineer — HPC & Automation (Silicon Engineering)
Redmond, Washington, United States
$125k-$175k/yr OnsiteFull Time
SpaceX
SpaceX: Designs and launches advanced rockets and satellite internet constellations.
2+ YOEBachelor's in CS/IS/engineering or 2+ years SRE/HPC/system administration experience; 1+ year dev experience with Bash/Python; 1+ year Linux experience; familiarity with containers, IaC, CI/CD, monitoring, databases, networking, and ASIC tool flows.
Bash, Python, Linux, Docker, Kubernetes, MySQL, PostgreSQL, SQLite, TCP/IP, Slurm, LSF, Terraform, Ansible, Puppet, Grafana, Prometheus, Jenkins, Bamboo, REST API, NetApp ONTAP REST API/CLI, Cadence, Synopsys, Ansys, Keysight, Siemens, Grok, Claude Code
4d
Save
Mark Applied
Hide
Staff Software Engineer, HPC
Foster City or Boston or Seattle
$230k-$295k/yr HybridFull Time
Zoox
ZooxNASDAQ: AMZN: Developing autonomous robotaxis for urban ride-hailing services.
Experience operating large-scale distributed systems, Ray.io, Kubernetes, AWS or similar cloud infrastructure, reliable scalable systems, cross-functional planning, and Python proficiency.
Ray.io, SLURM, Kubernetes, AWS, Python
2w
Save
Mark Applied
Hide
Security engineer, detection and response
San Francisco or Seattle or New York City
$132k-$258k/yr HybridFull Time
Writer
Writer: Platform for building and deploying enterprise generative AI agents.
3+ YOE3+ years in security operations, detection engineering, or incident response; experience securing AI/ML or distributed HPC infrastructure; strong programming and forensic skills; SIEM and detection experience.
Python, KQL, SPL, SIEM
4w
Save
Mark Applied
Hide
Senior AI Network Systems Engineer
Redmond or Mountain View or Hillsboro
$120k-$235k/yr HybridFull Time
Microsoft
MicrosoftNASDAQ: MSFT: Develops software, services, devices, and cloud computing solutions.
5+ YOEMaster's+3yrs or Bachelor's+5yrs in engineering fields; 5+ years networking for accelerator systems, Ethernet and RDMA fabrics; experience with AI/HPC/cloud deployments; security screening required.
SONiC, Linux, Azure
2mo
Save
Mark Applied
Hide
Senior Specialist Field Engineer - Compute Infrastructure
Livingston or New York or Sunnyvale or San Francisco or Bellevue or Dallas
$188k-$275k/yr FieldFull Time
CoreWeave
CoreWeaveNASDAQ: CRWV: Cloud platform providing GPU-accelerated infrastructure for AI workloads.
7+ YOEB.S. or equivalent experience, 7+ years in solutions architecture/field/infrastructure engineering or TAM for cloud/HPC; deep expertise with bare-metal GPU clusters, Linux, networking, InfiniBand/NVLink, PXE, and customer-facing technical leadership.
Kubernetes, Slurm, Python, Bash, Ansible, NCCL, ib_write_bw, InfiniBand, NVLink, NVIDIA HGX, GB200, Linux, PXE, BMC, BIOS, TCP/IP, Bare Metal as a Service (BMaaS)
2mo
Save
Mark Applied
Hide
Senior Network Engineer
New York City or San Francisco or Seattle or London or United States
$150k-$190k/yr RemoteFull Time
Lightning AI
Lightning AI: Unified platform to build, train, and deploy AI models.
5+ YOE5+ years large-scale data center networking experience with Cumulus NOS, spine/leaf design, BGP/EVPN/VXLAN, HPC/GPU networking, automation (Python/Ansible/Terraform), and strong documentation skills.
Cumulus NOS, SONiC, Junos, Python, Ansible, Terraform, EVPN, VXLAN, BGP, RoCE, RDMA, InfiniBand, VPC, NFV, Direct Connect, Cloud Connect, NVIDIA Spectrum, NVIDIA Quantum, NVIDIA BlueField
4w
Save
Mark Applied
Hide
Principal Core Infrastructure Engineer
Seattle, Washington, United States
$85k-$210k/yr OnsiteFull Time
Oracle
OracleNYSE: ORCL: Provides cloud infrastructure and enterprise software for global businesses.
6+ YOEDesign, deploy, and manage AI/ML and HPC infrastructure; scripting and automation (Ansible,Terraform,Python); containerization (Docker,Kubernetes); strong Linux, networking, security, and troubleshooting skills.
Ansible, Terraform, Python, Kubernetes, Docker, Slurm, PBS, TensorFlow, PyTorch, scikit-learn, Jenkins, GitLab CI/CD, Prometheus, GitHub, Scala, Oracle Linux, RHEL, CentOS, Ubuntu, Debian
1mo
Save
Mark Applied
Hide
ECAD Application Engineer, Design Technologies
Sunnyvale or Redmond
$117k-$185k/yr OnsiteFull Time
Amazon
AmazonNASDAQ: AMZN: Global online retail and cloud computing technology provider.
3+ YOEBachelor's in electrical engineering, 3+ years working with Cadence OrCAD/Allegro, System Capture, and PDM/PLM; experience integrating AI into workflows preferred; must be a U.S. citizen.
Cadence OrCAD, Allegro, System Capture, Pulse, Product Data Management (PDM), Product Lifecycle Management (PLM), High Performance Compute (HPC), CAD, EDA
16h
Save
Mark Applied
Hide
Senior Site Reliability Engineer
San Francisco or Bellevue
$240k-$356k/yr HybridFull Time
Lambda
Lambda: Provides high-performance GPU cloud infrastructure for AI development.
7+ YOERequires 7+ years in SRE, HPC engineering, DevOps, or similar; expertise in AI infrastructure, Linux, distributed systems, networking, Python, Go, monitoring, and automation tools.
Linux, Ansible, Terraform, Python, Go, Prometheus, Grafana, ClickHouse, PyTorch, TensorFlow, DeepSpeed, MLPerf, Docker, Kubernetes, InfiniBand, RoCE, NCCL, GPU-direct, CLOS, 100GbE, Ethernet, SOC 2, ISO 27001
3w
Save
Mark Applied
Hide
Software Engineer, Systems ML (Technical Leadership)
Bellevue, Washington, United States
$219k-$301k/yr OnsiteFull Time
Meta
MetaNASDAQ: META: Develops social networking platforms and virtual reality technologies.
12+ YOEBachelor's degree or equivalent,12+ years software engineering focused on ML systems (AI infrastructure, compilers, HPC, GPU, frameworks),experience delivering large-scale ML training/inference infrastructure,leadership of cross-functional initiatives,proficiency in C++,Python,or CUDA.
C++, Python, CUDA, MLIR, XLA, TVM
1mo
Save
Mark Applied
Hide
Senior Engineering Manager, AI Infrastructure
Seattle, Washington, United States
$147k-$220k/yr OnsiteFull Time
Allen Institute for Artificial Intelligence
Allen Institute for Artificial Intelligence: Non-profit institute developing open-source artificial intelligence for scientific impact.
12+ YOE2+ Mgmt12+ years infrastructure/HPC experience, 2+ years supervising engineering teams, Bachelor's in a related field, hands-on GPU/HPC, orchestration, storage, Go or Python proficiency.
Linux kernel, InfiniBand, NCCL, Beaker, AWS, GCP, Kubernetes, Slurm, WEKA, Ceph, Lustre, Go, Python
4w
Save
Mark Applied
Hide
Security Engineer, Cloud
Austin or New York City or San Francisco or Seattle or United States
$218k-$252k/yr OnsiteFull Time
Fluidstack
Fluidstack: Provides high-performance cloud GPU infrastructure for AI development.
Hands-on experience securing cloud and bare-metal infrastructure, building security tooling and guardrails as code, incident response, and strong coding skills for production systems.
Kubernetes, eBPF, HPC
1mo
Save
Mark Applied
Hide
Sr Tech Sales Engineer Rep PacWest-AZ
Morrisville or Seattle or Portland or Boise or Vancouver
$115k-$130k/yr OnsiteFull Time
Lenovo
LenovoHKSE: 992: Manufactures personal computers, mobile devices, and server infrastructure.
2+ YOEBachelor's or associate in CS/CE or equivalent experience; 2+ years technical sales (or 5+ years data center sysadmin); knowledge of virtualization, storage, cloud, Linux, and HPC; strong consultative selling and communication skills; must reside in Pacific Northwest/AZ territory.
VMware, Linux, Hyper-V, Windows Server, NetApp Storage, Azure, Nutanix, Veeam, Commvault, Slurm, PBS, LSF, Lustre, BeeGFS, IBM Storage Scale, MPI
4d
Save
Mark Applied
Hide
Snowflake Advanced AI Solution Lead
San Francisco or Milwaukee or Dallas or Columbus or Kirkland or Cincinnati or New York City or Cleveland or Oklahoma City or Austin or Albany or Chicago or St. Petersburg or Hartford or Pittsburgh or St. Louis or Miami or Sacramento or Raleigh or Minneapolis or Mountain View or Scottsdale or Morristown or Denver or Boston or Philadelphia or Des Moines or Overland Park or Los Angeles or Charlotte or Walnut Creek or Carmel or Seattle or Houston or Arlington or Atlanta or Redmond or Bentonville or Beaverton or Nashville or Detroit or San Diego
$80k-$294k/yr HybridFull Time
Accenture
AccentureNYSE: ACN: Global provider of management consulting and technology services.
5+ YOERequires 5+ years in machine learning engineering or science, 5+ years in computer science foundations, 3+ years in distributed systems, 2+ years with Snowflake Cortex AI and cloud AI/ML deployment, plus a bachelor's degree or equivalent experience.
Snowflake Cortex AI, Cortex Agents, Cortex Analyst, Cortex Search, Deep Learning, Generative AI, Large Language Models, DevOps, MLOps, LLMOps, HPC, Python, TensorFlow, PyTorch, LLM, C++, Java, R, SQL