61 hpc engineer jobs at 38 companies in Fairfield, CA

2mo
Save
Mark Applied
Hide
Staff HPC Engineer
San Francisco, California, United States
$214k-$300k/yr HybridFull Time
Chan Zuckerberg Biohub
Chan Zuckerberg Biohub: Builds AI-powered technologies for biomedical research and disease study.
10+ YOE10+ years HPC infra, hybrid on-prem/cloud, GPU AI workloads; Slurm, Kubernetes; PyTorch/TensorFlow/JAX; MLOps; strong collaboration and leadership.
Slurm, Kubernetes, SUNK, Docker, Singularity, Terraform, Ansible, Python, Bash, Git, Horovod, DeepSpeed, Ray, PyTorch, TensorFlow, JAX, RAPIDS, Coreweave, AWS, GCP
1mo
Save
Mark Applied
Hide
HPC Systems Engineer
San Francisco, California, United States
$120k-$196k/yr HybridFull Time
University of California, San Francisco
University of California, San Francisco: Public research university dedicated to health sciences and education.
6+ YOEBachelor's in CS/engineering plus 6+ years HPC experience (or 10+ yrs related); expert HPC infrastructure design; experience with parallel filesystems (GPFS, Lustre, Vast, DDN); security to NIST/HIPAA standards; SLURM/PBS, Warewulf/InfiniBand, automated testing, advanced scripting, and technical documentation skills.
GPFS, Lustre, Vast, DDN, SLURM, PBS, Warewulf, InfiniBand
3w
Save
Mark Applied
Hide
Staff HPC Network Engineer
San Francisco, California, United States
$224k-$284k/yr OnsiteFull Time
Atoms
Atoms: Building specialized industrial robots and physical AI systems.
Experience designing and scaling high-performance networks for HPC/GPU compute, hands-on switch management (Arista EOS, Cisco NX-OS, Nvidia Cumulus, SONiC), packet analysis, fiber optics knowledge, cloud networking, config management, automation scripting, and monitoring tools.
Arista EOS, Cisco NX-OS, Nvidia Cumulus, SONiC, tcpdump, Wireshark, AWS, GCP, Azure, Ansible, Salt, Python, Prometheus, Grafana, ELK, GitHub
3w
Save
Mark Applied
Hide
Staff HPC Network Engineer
San Francisco, California, United States
$224k-$284k/yr OnsiteFull Time
Atoms
Atoms: Building robotics and software to automate physical world industries.
Experience designing and scaling networks for HPC/GPU compute; hands-on with Arista EOS, Cisco NX-OS, Nvidia Cumulus, and/or SONiC; packet analysis (tcpdump, Wireshark); fiber optics to 800 Gbps; familiarity with AWS/GCP/Azure, Ansible/Salt, Python, Prometheus, Grafana, ELK, and GitHub.
Arista EOS, Cisco NX-OS, Nvidia Cumulus, SONiC, tcpdump, Wireshark, AWS, GCP, Azure, Ansible, Salt, Python, Prometheus, Grafana, ELK, GitHub
1mo
Save
Mark Applied
Hide
HPC Systems Engineer
San Francisco, California, United States
HybridFull Time
UCSF Health
UCSF Health: Academic medical center providing advanced patient care and research.
6+ YOEBachelor's in a related field and 6+ years experience with large-scale/HPC systems (or 10+ years related). Expert HPC infrastructure, parallel filesystems, security controls (NIST/HIPPA), SLURM/PBS, Warewulf/InfiniBand, EM automation and advanced scripting.
GPFS, Lustre, Vast, DDN, NIST 800-171, NIST 800-223, HIPPA, SLURM, PBS, Warewulf, InfiniBand
1mo
Save
Mark Applied
Hide
HPC Scientific Support Engineer
Berkeley, California, United States
$139k-$268k/yr HybridFull Time
Lawrence Berkeley National Laboratory
Lawrence Berkeley National Laboratory: Conducts multidisciplinary scientific research for the U.S. Department of Energy.
8+ YOE8+ years related experience with a Bachelor's (CSE3) or 12+ years (CSE4); 2+ years using HPC systems; deep HPC expertise (Linux/Unix, Fortran, C/C++, MPI, OpenMP, CUDA, Python, containers, debuggers, performance tools); strong communication and teaching skills.
Linux/Unix, Fortran, C/C++, MPI, OpenMP, OpenACC, CUDA, shell scripting, Python, parallel algorithms, programming models, AI models, debuggers, performance tools, containers, REST, JavaScript, SQL
4w
Save
Mark Applied
Hide
HPC Scientific Support Engineer
Berkeley, California, United States
$139k-$268k/yr HybridFull Time
Lawrence Berkeley National Laboratory
Lawrence Berkeley National Laboratory: Conducts scientific research to address global challenges.
8+ YOE8+ years related experience with a Bachelor's (CSE3) or 12+ years (CSE4); 2+ years using HPC systems; expertise in high-performance scientific computing including at least four of Linux/Unix, Fortran, C/C++, MPI, OpenMP, OpenACC, CUDA, shell scripting, Python, parallel algorithms, debuggers, performance tools, containers; strong communication and t...
Linux/Unix, Fortran, C/C++, MPI, OpenMP, OpenACC, CUDA, shell scripting, Python, debuggers, performance tools, containers, REST, JavaScript, SQL
1mo
Save
Mark Applied
Hide
HPC/ML Infrastructure Engineer
San Francisco or Tokyo
OnsiteFull Time
Spellbrush
Spellbrush: Develops anime-themed video games using proprietary generative AI technology.
Experienced HPC/ML infrastructure engineer with Linux sysadmin skills, cluster bring-up and operations experience, familiarity with SLURM and parallel filesystems, networking and datacenter hardware handling.
SLURM, Slinky, K8s, Warewulf, MAAS, Ansible, WEKA, VAST, Ceph, Tailscale, Grafana, Prometheus, LDAP, dmesg, HGX, VLAN
3w
Save
Mark Applied
Hide
GPU Systems Engineer – HPC / Parallel Computing
San Francisco or Los Angeles
$160k-$320k/yr OnsiteFull Time
Vast.ai
Vast.ai: Decentralized marketplace for GPU cloud computing resources.
Experience with HPC/parallel programming, GPU systems optimization, CUDA/C++, Python, and Linux; familiarity with parallel frameworks (HIP, SYCL, OpenCL, OpenACC) and HPC performance tooling.
CUDA, C++, GPGPU, Python, Linux, C++17, C++20, HIP, SYCL, OpenCL, OpenACC
2w
Save
Mark Applied
Hide
HPC Platform Engineer, Software, Center for Quantum Computing
San Francisco, California, United States
OnsiteFull Time
Amazon
AmazonNASDAQ: AMZN: Global online retail and cloud computing technology provider.
2+ YOEExperience automating and supporting large-scale infrastructure, programming in at least one modern language, Linux/Unix, CI/CD and infrastructure-as-code, and 2+ years designing or architecting systems.
Python, Ruby, Golang, Java, C++, C#, Rust, Linux, MPI, Docker, Kubernetes, AWS CDK, CloudFormation, EC2, S3, EBS, SQS, Lambda, VPC, DNS, DHCP, TCP/IP, HTTP, Palace
1mo
Save
Mark Applied
Hide
Staff Engineer, Distributed Storage and HPC & AI Infrastructure
San Francisco, California, United States
$250k-$300k/yr HybridFull Time
Together AI
Together AI: Cloud platform for training and deploying artificial intelligence models.
8+ YOE8+ years storage engineering experience with 3+ years managing multi-petabyte distributed storage; Kubernetes and cloud-native storage expertise; strong Go and Python skills; BS/MS or equivalent experience.
WekaFS, Ceph, Lustre, GPFS, BeeGFS, S3, MinIO, R2, Kubernetes, CSI, StatefulSets, PersistentVolumes, Go, Python, RDMA, InfiniBand, NVMe-oF, iSCSI, Helm, Terraform, Ansible, ArgoCD, ext4, xfs, LVM, NVMe, RAID, Prometheus, Grafana, Thanos, GPU Direct Storage (GDS), kubebuilder, controller-runtime, Velero, Restic, fio, iperf3, iostat, blktrace, GitOps
3mo
Save
Mark Applied
Hide
Senior Aerodynamics Engineer
San Francisco, California, United States
$160k-$215k/yr OnsiteFull Time
Astro Mechanica
Astro Mechanica: Building turboelectric jet engines for supersonic aircraft.
10+ YOEBachelor’s in Aerospace or Mechanical Engineering; 10+ years in aerodynamic design, CFD, and optimization; proficiency with CFD solvers, parametric geometry, wind tunnel testing; experience with adjoint-based optimization; strong multidisciplinary data synthesis; HPC experience; leadership ability.
CFD, parametric geometry modeling, wind tunnel testing, adjoint-based optimization, high performance computing (HPC), AWS, Azure, Google HPC
2mo
Save
Mark Applied
Hide
ML Research Engineer - Research
San Francisco or New York City
HybridFull Time
Achira
Achira: AI foundation models for atomistic simulation and drug discovery.
Strong software engineering, ML research experience, PyTorch/JAX, HPC, and PhD or equivalent work in ML/scientific computing.
PyTorch, JAX, Git, Python, HPC
2d
Save
Mark Applied
Hide
Principal Engineer, CAPE
San Francisco, California, United States
$285k-$335k/yr OnsiteFull Time
Crusoe
Crusoe: Provides energy-efficient cloud infrastructure powered by stranded and renewable energy.
10+ YOE10+ years building infrastructure-layer systems, distributed systems and control-plane experience, GPU/HPC infrastructure fluency, systems language expertise (Go/Rust/C++), and experience with observability and failure prediction.
Go, Rust, C++, NVLink, InfiniBand, RoCE
2mo
Save
Mark Applied
Hide
Design Optimization Engineer - Postdoctoral Researcher
Livermore, California, United States
$122k-$143k/yr HybridFull Time
Lawrence Livermore National Laboratory
Lawrence Livermore National Laboratory: Develops science and technology for United States national security.
PhD in Engineering/Mathematics/Computational Science; HPC experience; C/C++/FORTRAN; Python; MPI/OpenMP/CUDA; ability to perform independent research.
C, C++, FORTRAN, Python, MPI, OpenMP, CUDA, Finite Elements, High-Performance Computing
1mo
Save
Mark Applied
Hide
Senior Specialist Field Engineer - Compute Infrastructure
Livingston or New York or Sunnyvale or San Francisco or Bellevue or Dallas
$188k-$275k/yr FieldFull Time
CoreWeave
CoreWeaveNASDAQ: CRWV: Cloud platform providing GPU-accelerated infrastructure for AI workloads.
7+ YOEB.S. or equivalent experience, 7+ years in solutions architecture/field/infrastructure engineering or TAM for cloud/HPC; deep expertise with bare-metal GPU clusters, Linux, networking, InfiniBand/NVLink, PXE, and customer-facing technical leadership.
Kubernetes, Slurm, Python, Bash, Ansible, NCCL, ib_write_bw, InfiniBand, NVLink, NVIDIA HGX, GB200, Linux, PXE, BMC, BIOS, TCP/IP, Bare Metal as a Service (BMaaS)
1mo
Save
Mark Applied
Hide
Senior Network Engineer
New York City or San Francisco or Seattle or London or United States
$150k-$190k/yr RemoteFull Time
Lightning AI
Lightning AI: Unified platform to build, train, and deploy AI models.
5+ YOE5+ years large-scale data center networking experience with Cumulus NOS, spine/leaf design, BGP/EVPN/VXLAN, HPC/GPU networking, automation (Python/Ansible/Terraform), and strong documentation skills.
Cumulus NOS, SONiC, Junos, Python, Ansible, Terraform, EVPN, VXLAN, BGP, RoCE, RDMA, InfiniBand, VPC, NFV, Direct Connect, Cloud Connect, NVIDIA Spectrum, NVIDIA Quantum, NVIDIA BlueField
1mo
Save
Mark Applied
Hide
Data Center Engineer II
San Francisco, California, United States
$169k-$232k/yr OnsiteFull Time
Adyen
AdyenEuronext Amsterdam: ADYEN: Unified payment platform for global business commerce.
Experienced data center engineer with DCIM and asset management expertise, proven infrastructure design and delivery, strong project management, hands-on field operations, and willingness to travel internationally (20-30%).
DCIM, High-Performance Computing (HPC)
2mo
Save
Mark Applied
Hide
Associate Director, Scientific Computing and AI Engineer
South San Francisco, California, United States
$159k-$207k/yr OnsiteAll Commitments Available
Denali Therapeutics
Denali TherapeuticsNASDAQ: DNLI: Developing therapies to treat neurodegenerative and lysosomal storage diseases.
10+ YOE10-12 years in platform engineering/infrastructure with senior IC; expertise in HPC, cloud, AI/ML; IaC/CI/CD; Python; regulated biotech experience; leadership experience.
Slurm, LSF, GPU/CUDA, Cloud platforms, MLOps, LangChain, LangGraph, AutoGen, Python, Terraform, CI/CD
2w
Save
Mark Applied
Hide
Network Operations Engineer, AI Networking
San Francisco, California, United States
$157k-$221k/yr OnsiteFull Time
OpenAI
OpenAI: Develops artificial intelligence models and generative AI software services.
5+ YOE5+ years operating large-scale data center, cloud, AI, or HPC networks; strong L2/L3 networking knowledge; hands-on with switches, optics, and routing; Python automation and observability experience.
Cisco NX-OS, Arista EOS, NVIDIA Spectrum, Cumulus Linux, Juniper JunOS, Prometheus, Grafana, gNMI, SNMP, Python, Git, REST APIs, Terraform, AWS, Azure, Google Cloud, VAST, DDN, NVIDIA HGX, DGX, GB200