UCSF Health: Academic medical center providing advanced patient care and research.
6+ YOEBachelor's in a related field and 6+ years experience with large-scale/HPC systems (or 10+ years related). Expert HPC infrastructure, parallel filesystems, security controls (NIST/HIPPA), SLURM/PBS, Warewulf/InfiniBand, EM automation and advanced scripting.
Lawrence Berkeley National Laboratory: Conducts multidisciplinary scientific research for the U.S. Department of Energy.
8+ YOE8+ years related experience with a Bachelor's (CSE3) or 12+ years (CSE4); 2+ years using HPC systems; deep HPC expertise (Linux/Unix, Fortran, C/C++, MPI, OpenMP, CUDA, Python, containers, debuggers, performance tools); strong communication and teaching skills.
Lawrence Berkeley National Laboratory: Conducts scientific research to address global challenges.
8+ YOE8+ years related experience with a Bachelor's (CSE3) or 12+ years (CSE4); 2+ years using HPC systems; expertise in high-performance scientific computing including at least four of Linux/Unix, Fortran, C/C++, MPI, OpenMP, OpenACC, CUDA, shell scripting, Python, parallel algorithms, debuggers, performance tools, containers; strong communication and t...
Spellbrush: Develops anime-themed video games using proprietary generative AI technology.
Experienced HPC/ML infrastructure engineer with Linux sysadmin skills, cluster bring-up and operations experience, familiarity with SLURM and parallel filesystems, networking and datacenter hardware handling.
Vast.ai: Decentralized marketplace for GPU cloud computing resources.
Experience with HPC/parallel programming, GPU systems optimization, CUDA/C++, Python, and Linux; familiarity with parallel frameworks (HIP, SYCL, OpenCL, OpenACC) and HPC performance tooling.
HPC Platform Engineer, Software, Center for Quantum Computing
San Francisco, California, United States
OnsiteFull Time
AmazonNASDAQ: AMZN: Global online retail and cloud computing technology provider.
2+ YOEExperience automating and supporting large-scale infrastructure, programming in at least one modern language, Linux/Unix, CI/CD and infrastructure-as-code, and 2+ years designing or architecting systems.
Staff Engineer, Distributed Storage and HPC & AI Infrastructure
San Francisco, California, United States
$250k-$300k/yrHybridFull Time
Together AI: Cloud platform for training and deploying artificial intelligence models.
8+ YOE8+ years storage engineering experience with 3+ years managing multi-petabyte distributed storage; Kubernetes and cloud-native storage expertise; strong Go and Python skills; BS/MS or equivalent experience.
Astro Mechanica: Building turboelectric jet engines for supersonic aircraft.
10+ YOEBachelor’s in Aerospace or Mechanical Engineering; 10+ years in aerodynamic design, CFD, and optimization; proficiency with CFD solvers, parametric geometry, wind tunnel testing; experience with adjoint-based optimization; strong multidisciplinary data synthesis; HPC experience; leadership ability.
CFD, parametric geometry modeling, wind tunnel testing, adjoint-based optimization, high performance computing (HPC), AWS, Azure, Google HPC
Crusoe: Provides energy-efficient cloud infrastructure powered by stranded and renewable energy.
10+ YOE10+ years building infrastructure-layer systems, distributed systems and control-plane experience, GPU/HPC infrastructure fluency, systems language expertise (Go/Rust/C++), and experience with observability and failure prediction.
Senior Specialist Field Engineer - Compute Infrastructure
Livingston or New York or Sunnyvale or San Francisco or Bellevue or Dallas
$188k-$275k/yrFieldFull Time
CoreWeaveNASDAQ: CRWV: Cloud platform providing GPU-accelerated infrastructure for AI workloads.
7+ YOEB.S. or equivalent experience, 7+ years in solutions architecture/field/infrastructure engineering or TAM for cloud/HPC; deep expertise with bare-metal GPU clusters, Linux, networking, InfiniBand/NVLink, PXE, and customer-facing technical leadership.
Kubernetes, Slurm, Python, Bash, Ansible, NCCL, ib_write_bw, InfiniBand, NVLink, NVIDIA HGX, GB200, Linux, PXE, BMC, BIOS, TCP/IP, Bare Metal as a Service (BMaaS)
New York City or San Francisco or Seattle or London or United States
$150k-$190k/yrRemoteFull Time
Lightning AI: Unified platform to build, train, and deploy AI models.
5+ YOE5+ years large-scale data center networking experience with Cumulus NOS, spine/leaf design, BGP/EVPN/VXLAN, HPC/GPU networking, automation (Python/Ansible/Terraform), and strong documentation skills.
AdyenEuronext Amsterdam: ADYEN: Unified payment platform for global business commerce.
Experienced data center engineer with DCIM and asset management expertise, proven infrastructure design and delivery, strong project management, hands-on field operations, and willingness to travel internationally (20-30%).
OpenAI: Develops artificial intelligence models and generative AI software services.
5+ YOE5+ years operating large-scale data center, cloud, AI, or HPC networks; strong L2/L3 networking knowledge; hands-on with switches, optics, and routing; Python automation and observability experience.