Site Reliability Engineer — HPC & Automation (Silicon Engineering)
Redmond, Washington, United States
$125k-$175k/yrOnsiteFull Time
SpaceX: Designs and launches advanced rockets and satellite internet constellations.
2+ YOEBachelor's in CS/IS/engineering or 2+ years SRE/HPC/system administration experience; 1+ year dev experience with Bash/Python; 1+ year Linux experience; familiarity with containers, IaC, CI/CD, monitoring, databases, networking, and ASIC tool flows.
15+ YOEBachelor's in CS or equivalent, 15+ years software engineering experience with distributed systems/ML/AI infrastructure, expertise in large-scale systems, deep learning frameworks, and influencing technical direction.
MicrosoftNASDAQ: MSFT: Develops software, services, devices, and cloud computing solutions.
5+ YOEMaster's+3yrs or Bachelor's+5yrs in engineering fields; 5+ years networking for accelerator systems, Ethernet and RDMA fabrics; experience with AI/HPC/cloud deployments; security screening required.
Senior Specialist Field Engineer - Compute Infrastructure
Livingston or New York or Sunnyvale or San Francisco or Bellevue or Dallas
$188k-$275k/yrFieldFull Time
CoreWeaveNASDAQ: CRWV: Cloud platform providing GPU-accelerated infrastructure for AI workloads.
7+ YOEB.S. or equivalent experience, 7+ years in solutions architecture/field/infrastructure engineering or TAM for cloud/HPC; deep expertise with bare-metal GPU clusters, Linux, networking, InfiniBand/NVLink, PXE, and customer-facing technical leadership.
Kubernetes, Slurm, Python, Bash, Ansible, NCCL, ib_write_bw, InfiniBand, NVLink, NVIDIA HGX, GB200, Linux, PXE, BMC, BIOS, TCP/IP, Bare Metal as a Service (BMaaS)
New York City or San Francisco or Seattle or London or United States
$150k-$190k/yrRemoteFull Time
Lightning AI: Unified platform to build, train, and deploy AI models.
5+ YOE5+ years large-scale data center networking experience with Cumulus NOS, spine/leaf design, BGP/EVPN/VXLAN, HPC/GPU networking, automation (Python/Ansible/Terraform), and strong documentation skills.
San Francisco or New York City or Seattle or London
$230k-$405k/yrHybridFull Time
OpenAI: Develops artificial intelligence models and generative AI software services.
Strong software engineering skills; experience with production infrastructure; distributed systems or HPC experience; ability to debug complex systems and improve reliability.
Software Engineer, Systems ML (Technical Leadership)
Bellevue, Washington, United States
$219k-$301k/yrOnsiteFull Time
MetaNASDAQ: META: Develops social networking platforms and virtual reality technologies.
12+ YOEBachelor's degree or equivalent,12+ years software engineering focused on ML systems (AI infrastructure, compilers, HPC, GPU, frameworks),experience delivering large-scale ML training/inference infrastructure,leadership of cross-functional initiatives,proficiency in C++,Python,or CUDA.
ZooxNASDAQ: AMZN: Developing autonomous robotaxis for urban ride-hailing services.
8+ YOEBachelor's degree and 8+ years experience; 4+ years Bazel (or Buck/Pants); Linux and Python proficiency; experience with large polyglot codebases, HPC/compute infrastructure, and cost/efficiency optimization.
Bazel, Buck, Pants, Linux, Python, Ray.io, SLURM, Kubernetes, Starlark, C++
Allen Institute for Artificial Intelligence: Non-profit institute developing open-source artificial intelligence for scientific impact.
12+ YOE2+ Mgmt12+ years infrastructure/HPC experience, 2+ years supervising engineering teams, Bachelor's in a related field, hands-on GPU/HPC, orchestration, storage, Go or Python proficiency.
Austin or New York City or San Francisco or Seattle or United States
$218k-$252k/yrOnsiteFull Time
Fluidstack: Provides high-performance cloud GPU infrastructure for AI development.
Hands-on experience securing cloud and bare-metal infrastructure, building security tooling and guardrails as code, incident response, and strong coding skills for production systems.
Senior Software Engineer, DGX Cloud AI Infrastructure
Santa Clara or Austin or Redmond or Oregon or Washington
$184k-$357k/yrHybridFull Time
NVIDIANASDAQ: NVDA: Designs GPU-accelerated computing and artificial intelligence hardware.
8+ YOE8+ years building software infrastructure for large-scale AI/HPC, expertise debugging multi-GPU and multi-node workloads, NCCL/CUDA experience, expert Python and C/C++, and experience with containerized cluster environments.
Morrisville or Seattle or Portland or Boise or Vancouver
$115k-$130k/yrOnsiteFull Time
LenovoHKSE: 992: Manufactures personal computers, mobile devices, and server infrastructure.
2+ YOEBachelor's or associate in CS/CE or equivalent experience; 2+ years technical sales (or 5+ years data center sysadmin); knowledge of virtualization, storage, cloud, Linux, and HPC; strong consultative selling and communication skills; must reside in Pacific Northwest/AZ territory.
VMware, Linux, Hyper-V, Windows Server, NetApp Storage, Azure, Nutanix, Veeam, Commvault, Slurm, PBS, LSF, Lustre, BeeGFS, IBM Storage Scale, MPI
Pacific Northwest National Laboratory: Department of Energy research laboratory focused on scientific innovation.
2+ YOEDegree in computer science, engineering, mathematics, or related field; 2+ years experience or higher degree; strong ML, PyTorch, and HPC skills.