Senior HPC Support Engineer - Compute and GPU Platform
California or North Carolina or New York or Washington or Massachusetts or Santa Clara
$108k-$207k/yrRemoteFull Time
NVIDIANASDAQ: NVDA: Designs GPU-accelerated computing and artificial intelligence hardware.
5+ YOE5+ years supporting and debugging hardware/software; strong Linux system administration, server and datacenter knowledge; experience with containers, virtualization, and AI technologies; degree in networking/CS/EE or equivalent experience; excellent English.
Linux, Red Hat Enterprise Linux, Ubuntu, Microsoft Windows, VMware, Docker, Kubernetes, Bash, Python, DGX Platform, InfiniBand, RDMA, RoCEv2, MPI, NCCL
Site Reliability Engineer — HPC & Automation (Silicon Engineering)
Redmond, Washington, United States
$125k-$175k/yrOnsiteFull Time
SpaceX: Designs and launches advanced rockets and satellite internet constellations.
2+ YOEBachelor's in CS/IS/engineering or 2+ years SRE/HPC/system administration experience; 1+ year dev experience with Bash/Python; 1+ year Linux experience; familiarity with containers, IaC, CI/CD, monitoring, databases, networking, and ASIC tool flows.
Lynchburg or Virginia or Washington or North Carolina or Pennsylvania or Massachusetts
$100k-$110k/yrOnsiteFull Time
Framatome: Designs and services nuclear power reactors and fuel systems.
3+ YOEBachelor's in Engineering (or related), minimum 3 years related experience, neutronics and fuel modeling experience, familiarity with cross-section, 3D simulator and Monte Carlo codes, programming (FORTRAN, PYTHON, C++), Linux and HPC.
15+ YOEBachelor's in CS or equivalent, 15+ years software engineering experience with distributed systems/ML/AI infrastructure, expertise in large-scale systems, deep learning frameworks, and influencing technical direction.
MicrosoftNASDAQ: MSFT: Develops software, services, devices, and cloud computing solutions.
5+ YOEMaster's+3yrs or Bachelor's+5yrs in engineering fields; 5+ years networking for accelerator systems, Ethernet and RDMA fabrics; experience with AI/HPC/cloud deployments; security screening required.
Senior Specialist Field Engineer - Compute Infrastructure
Livingston or New York or Sunnyvale or San Francisco or Bellevue or Dallas
$188k-$275k/yrFieldFull Time
CoreWeaveNASDAQ: CRWV: Cloud platform providing GPU-accelerated infrastructure for AI workloads.
7+ YOEB.S. or equivalent experience, 7+ years in solutions architecture/field/infrastructure engineering or TAM for cloud/HPC; deep expertise with bare-metal GPU clusters, Linux, networking, InfiniBand/NVLink, PXE, and customer-facing technical leadership.
Kubernetes, Slurm, Python, Bash, Ansible, NCCL, ib_write_bw, InfiniBand, NVLink, NVIDIA HGX, GB200, Linux, PXE, BMC, BIOS, TCP/IP, Bare Metal as a Service (BMaaS)
New York City or San Francisco or Seattle or London or United States
$150k-$190k/yrRemoteFull Time
Lightning AI: Unified platform to build, train, and deploy AI models.
5+ YOE5+ years large-scale data center networking experience with Cumulus NOS, spine/leaf design, BGP/EVPN/VXLAN, HPC/GPU networking, automation (Python/Ansible/Terraform), and strong documentation skills.
San Francisco or New York City or Seattle or London
$230k-$405k/yrHybridFull Time
OpenAI: Develops artificial intelligence models and generative AI software services.
Strong software engineering skills; experience with production infrastructure; distributed systems or HPC experience; ability to debug complex systems and improve reliability.
Software Engineer, Systems ML (Technical Leadership)
Bellevue, Washington, United States
$219k-$301k/yrOnsiteFull Time
MetaNASDAQ: META: Develops social networking platforms and virtual reality technologies.
12+ YOEBachelor's degree or equivalent,12+ years software engineering focused on ML systems (AI infrastructure, compilers, HPC, GPU, frameworks),experience delivering large-scale ML training/inference infrastructure,leadership of cross-functional initiatives,proficiency in C++,Python,or CUDA.
ZooxNASDAQ: AMZN: Developing autonomous robotaxis for urban ride-hailing services.
8+ YOEBachelor's degree and 8+ years experience; 4+ years Bazel (or Buck/Pants); Linux and Python proficiency; experience with large polyglot codebases, HPC/compute infrastructure, and cost/efficiency optimization.
Bazel, Buck, Pants, Linux, Python, Ray.io, SLURM, Kubernetes, Starlark, C++
Allen Institute for Artificial Intelligence: Non-profit institute developing open-source artificial intelligence for scientific impact.
12+ YOE2+ Mgmt12+ years infrastructure/HPC experience, 2+ years supervising engineering teams, Bachelor's in a related field, hands-on GPU/HPC, orchestration, storage, Go or Python proficiency.
Austin or New York City or San Francisco or Seattle or United States
$218k-$252k/yrOnsiteFull Time
Fluidstack: Provides high-performance cloud GPU infrastructure for AI development.
Hands-on experience securing cloud and bare-metal infrastructure, building security tooling and guardrails as code, incident response, and strong coding skills for production systems.
Morrisville or Seattle or Portland or Boise or Vancouver
$115k-$130k/yrOnsiteFull Time
LenovoHKSE: 992: Manufactures personal computers, mobile devices, and server infrastructure.
2+ YOEBachelor's or associate in CS/CE or equivalent experience; 2+ years technical sales (or 5+ years data center sysadmin); knowledge of virtualization, storage, cloud, Linux, and HPC; strong consultative selling and communication skills; must reside in Pacific Northwest/AZ territory.
VMware, Linux, Hyper-V, Windows Server, NetApp Storage, Azure, Nutanix, Veeam, Commvault, Slurm, PBS, LSF, Lustre, BeeGFS, IBM Storage Scale, MPI
Pacific Northwest National Laboratory: Department of Energy research laboratory focused on scientific innovation.
2+ YOEDegree in computer science, engineering, mathematics, or related field; 2+ years experience or higher degree; strong ML, PyTorch, and HPC skills.