101 hpc engineer jobs at 59 companies in San Rafael, CA

2mo
Save
Mark Applied
Hide
HPC Engineer
Sunnyvale, California, United States
$150k-$300k/yr OnsiteFull Time
Institute of Foundation Models
Institute of Foundation Models: Develops open-source frontier-class AI foundation models and research.
2+ YOE2+ years in Linux systems administration, SRE, DevOps, cloud operations, HPC or infrastructure operations; strong Linux troubleshooting; scripting in Python or Bash; Bachelor’s in a related field.
Slurm, GPU infrastructure, AWS, Azure, GCP, Grafana, Prometheus, Datadog, Containers, Kubernetes, Python, Bash, Linux
2mo
Save
Mark Applied
Hide
HPC Systems Engineer
San Francisco, California, United States
$120k-$196k/yr HybridFull Time
University of California, San Francisco
University of California, San Francisco: Public research university dedicated to health sciences and education.
6+ YOEBachelor's in CS/engineering plus 6+ years HPC experience (or 10+ yrs related); expert HPC infrastructure design; experience with parallel filesystems (GPFS, Lustre, Vast, DDN); security to NIST/HIPAA standards; SLURM/PBS, Warewulf/InfiniBand, automated testing, advanced scripting, and technical documentation skills.
GPFS, Lustre, Vast, DDN, SLURM, PBS, Warewulf, InfiniBand
1mo
Save
Mark Applied
Hide
Staff HPC Network Engineer
San Francisco, California, United States
$224k-$284k/yr OnsiteFull Time
Atoms
Atoms: Building specialized industrial robots and physical AI systems.
Experience designing and scaling high-performance networks for HPC/GPU compute, hands-on switch management (Arista EOS, Cisco NX-OS, Nvidia Cumulus, SONiC), packet analysis, fiber optics knowledge, cloud networking, config management, automation scripting, and monitoring tools.
Arista EOS, Cisco NX-OS, Nvidia Cumulus, SONiC, tcpdump, Wireshark, AWS, GCP, Azure, Ansible, Salt, Python, Prometheus, Grafana, ELK, GitHub
1mo
Save
Mark Applied
Hide
Staff HPC Network Engineer
San Francisco, California, United States
$224k-$284k/yr OnsiteFull Time
Atoms
Atoms: Building robotics and software to automate physical world industries.
Experience designing and scaling networks for HPC/GPU compute; hands-on with Arista EOS, Cisco NX-OS, Nvidia Cumulus, and/or SONiC; packet analysis (tcpdump, Wireshark); fiber optics to 800 Gbps; familiarity with AWS/GCP/Azure, Ansible/Salt, Python, Prometheus, Grafana, ELK, and GitHub.
Arista EOS, Cisco NX-OS, Nvidia Cumulus, SONiC, tcpdump, Wireshark, AWS, GCP, Azure, Ansible, Salt, Python, Prometheus, Grafana, ELK, GitHub
2mo
Save
Mark Applied
Hide
HPC Scientific Support Engineer
Berkeley, California, United States
$139k-$268k/yr HybridFull Time
Lawrence Berkeley National Laboratory
Lawrence Berkeley National Laboratory: Conducts multidisciplinary scientific research for the U.S. Department of Energy.
8+ YOE8+ years related experience with a Bachelor's (CSE3) or 12+ years (CSE4); 2+ years using HPC systems; deep HPC expertise (Linux/Unix, Fortran, C/C++, MPI, OpenMP, CUDA, Python, containers, debuggers, performance tools); strong communication and teaching skills.
Linux/Unix, Fortran, C/C++, MPI, OpenMP, OpenACC, CUDA, shell scripting, Python, parallel algorithms, programming models, AI models, debuggers, performance tools, containers, REST, JavaScript, SQL
1mo
Save
Mark Applied
Hide
HPC Scientific Support Engineer
Berkeley, California, United States
$139k-$268k/yr HybridFull Time
Lawrence Berkeley National Laboratory
Lawrence Berkeley National Laboratory: Conducts scientific research to address global challenges.
8+ YOE8+ years related experience with a Bachelor's (CSE3) or 12+ years (CSE4); 2+ years using HPC systems; expertise in high-performance scientific computing including at least four of Linux/Unix, Fortran, C/C++, MPI, OpenMP, OpenACC, CUDA, shell scripting, Python, parallel algorithms, debuggers, performance tools, containers; strong communication and t...
Linux/Unix, Fortran, C/C++, MPI, OpenMP, OpenACC, CUDA, shell scripting, Python, debuggers, performance tools, containers, REST, JavaScript, SQL
2mo
Save
Mark Applied
Hide
HPC/ML Infrastructure Engineer
San Francisco or Tokyo
OnsiteFull Time
Spellbrush
Spellbrush: Develops anime-themed video games using proprietary generative AI technology.
Experienced HPC/ML infrastructure engineer with Linux sysadmin skills, cluster bring-up and operations experience, familiarity with SLURM and parallel filesystems, networking and datacenter hardware handling.
SLURM, Slinky, K8s, Warewulf, MAAS, Ansible, WEKA, VAST, Ceph, Tailscale, Grafana, Prometheus, LDAP, dmesg, HGX, VLAN
3w
Save
Mark Applied
Hide
Senior HPC Systems Architect
San Jose or San Francisco
$255k-$340k/yr HybridFull Time
Lambda
Lambda: Provides high-performance GPU cloud infrastructure for AI development.
8+ YOE8+ years designing large-scale HPC infrastructures, expertise with GPU clusters, high-speed networking, liquid cooling, benchmarking, capacity planning, and architecture documentation.
InfiniBand, Ethernet, Ansible, Terraform, Kubernetes
4w
Save
Mark Applied
Hide
AI/HPC Network Performance Engineer
Menlo Park, California, United States
$184k-$257k/yr OnsiteFull Time
Meta
MetaNASDAQ: META: Develops social networking platforms and virtual reality technologies.
8+ YOEBachelor's or equivalent,8+ years in system or network performance engineering for large-scale distributed/HPC environments; experience with datacenter networks, network automation, and coding in Python,C++,Go.
Python, C++, Go, IB, RDMA, RoCE
1w
Save
Mark Applied
Hide
Director, HPC Systems Software Engineering
Houston or New York City or San Francisco or Seattle
$230k-$343k/yr OnsiteFull Time
Nscale
Nscale: Vertically integrated AI infrastructure provider for high-performance computing.
Requires leadership of HPC or infrastructure engineering organizations, Linux, distributed systems, production operations, GPU infrastructure, networking, budgeting, and people management; Slurm experience strongly preferred.
Slurm, Kueue, Linux
1mo
Save
Mark Applied
Hide
GPU Systems Engineer – HPC / Parallel Computing
San Francisco or Los Angeles
$160k-$320k/yr OnsiteFull Time
Vast.ai
Vast.ai: Decentralized marketplace for GPU cloud computing resources.
Experience with HPC/parallel programming, GPU systems optimization, CUDA/C++, Python, and Linux; familiarity with parallel frameworks (HIP, SYCL, OpenCL, OpenACC) and HPC performance tooling.
CUDA, C++, GPGPU, Python, Linux, C++17, C++20, HIP, SYCL, OpenCL, OpenACC
4d
Save
Mark Applied
Hide
AI & HPC Infrastructure Engineer
Albany or Arlington or Atlanta or Austin or Beaverton or Bentonville or Boston or Carmel or Charlotte or Chicago or Cincinnati or Cleveland or Columbus or Culver City or Denver or Des Moines or Detroit or Hartford or Houston or Irvine or Irving or Kirkland or Miami or Milwaukee or Minneapolis or Morristown or Mountain View or Nashville or New York City or Oklahoma City or Overland Park or Philadelphia or Pittsburgh or Raleigh or Redmond or Sacramento or San Diego or San Francisco or Scottsdale or Seattle or St. Louis or St. Petersburg or Walnut Creek or United States
$80k-$266k/yr HybridFull Time
Accenture
AccentureNYSE: ACN: Global professional services firm providing consulting and technology solutions.
5+ YOERequires 5+ years designing AI infrastructure, accelerated computing, clusters, orchestration, and automation; bachelor's degree or equivalent experience. Strong Kubernetes, Python, Terraform, and cloud expertise required.
Slurm, Run:ai, Kubernetes, NVIDIA Base Command Manager (BCM), NVIDIA NGC, NCCL, NVLink, CUDA, TensorRT-LLM, vLLM, SGLang, Triton Inference Server, NVIDIA Dynamo, llm-d, MLPerf, NCCL tests, fio, iperf, InfiniBand, Ethernet, SONiC, NVMe, NVMe-oF, VAST, Weka, DDN, AWS, Azure, GCP, VMware, Nutanix, Python, Terraform, Ansible, REST, OpenAPI, JSON, YAML, TensorFlow, PyTorch, JAX, Jupyter notebooks, Google Colab, MCP
1mo
Save
Mark Applied
Hide
HPC Platform Engineer, Software, Center for Quantum Computing
San Francisco, California, United States
OnsiteFull Time
Amazon
AmazonNASDAQ: AMZN: Global online retail and cloud computing technology provider.
2+ YOEExperience automating and supporting large-scale infrastructure, programming in at least one modern language, Linux/Unix, CI/CD and infrastructure-as-code, and 2+ years designing or architecting systems.
Python, Ruby, Golang, Java, C++, C#, Rust, Linux, MPI, Docker, Kubernetes, AWS CDK, CloudFormation, EC2, S3, EBS, SQS, Lambda, VPC, DNS, DHCP, TCP/IP, HTTP, Palace
2mo
Save
Mark Applied
Hide
Staff Engineer, Distributed Storage and HPC & AI Infrastructure
San Francisco, California, United States
$250k-$300k/yr HybridFull Time
Together AI
Together AI: Cloud platform for training and deploying artificial intelligence models.
8+ YOE8+ years storage engineering experience with 3+ years managing multi-petabyte distributed storage; Kubernetes and cloud-native storage expertise; strong Go and Python skills; BS/MS or equivalent experience.
WekaFS, Ceph, Lustre, GPFS, BeeGFS, S3, MinIO, R2, Kubernetes, CSI, StatefulSets, PersistentVolumes, Go, Python, RDMA, InfiniBand, NVMe-oF, iSCSI, Helm, Terraform, Ansible, ArgoCD, ext4, xfs, LVM, NVMe, RAID, Prometheus, Grafana, Thanos, GPU Direct Storage (GDS), kubebuilder, controller-runtime, Velero, Restic, fio, iperf3, iostat, blktrace, GitOps
1mo
Save
Mark Applied
Hide
Staff Engineer, Infrastructure Platforms
Menlo Park, California, United States
$148k-$222k/yr HybridFull Time
PacBio
PacBioNASDAQ: PACB: Develops high-fidelity DNA sequencing systems for genomic research.
10+ YOE10+ years enterprise infrastructure engineering with Linux, hybrid cloud, storage, virtualization, HPC, IaC, automation; strong scripting (Python, Bash, PowerShell); incident response and security experience; US work authorization required.
Python, Bash, PowerShell, Kubernetes, GitOps, HPC
1w
Save
Mark Applied
Hide
Staff Software Engineer, HPC
Foster City or Boston or Seattle
$230k-$295k/yr HybridFull Time
Zoox
ZooxNASDAQ: AMZN: Developing autonomous robotaxis for urban ride-hailing services.
Experience operating large-scale distributed systems, Ray.io, Kubernetes, AWS or similar cloud infrastructure, reliable scalable systems, cross-functional planning, and Python proficiency.
Ray.io, SLURM, Kubernetes, AWS, Python
1mo
Save
Mark Applied
Hide
Staff Technical Product Manager – Electronics AI & HPC Simulation
Sunnyvale, California, United States
$117k-$175k/yr OnsiteFull Time
Synopsys
SynopsysNasdaq: SNPS: Provides software and IP for semiconductor design and manufacturing.
Bachelor's in engineering or CS preferred, strong Python scripting, experience with simulation/HPC/AI workflows and APIs, familiarity with HFSS/Icepak/Maxwell/AEDT, and experience with GPU/cloud HPC deployment.
Python, HFSS, Icepak, Maxwell, AEDT, NVIDIA Omniverse, APIs, HPC, GPU
2w
Save
Mark Applied
Hide
Security engineer, detection and response
San Francisco or Seattle or New York City
$132k-$258k/yr HybridFull Time
Writer
Writer: Platform for building and deploying enterprise generative AI agents.
3+ YOE3+ years in security operations, detection engineering, or incident response; experience securing AI/ML or distributed HPC infrastructure; strong programming and forensic skills; SIEM and detection experience.
Python, KQL, SPL, SIEM
1w
Save
Mark Applied
Hide
AI Infrastructure Engineer
San Francisco, California, United States
$150k-$220k/yr OnsiteFull Time
Sciforium
Sciforium: Building multimodal AI models and high-performance model serving infrastructure.
5+ YOE5+ years in systems or infrastructure engineering with GPU, HPC, or ML infrastructure experience; technical bachelor's or master's degree; Linux, Kubernetes, schedulers, configuration management, Python, Bash, containers, GPUs, and RDMA expertise.
Ansible, SaltStack, Git, Python, Bash, Kubernetes, NVIDIA GPU Operator, Slurm, Run:AI, enroot, pyxis, Docker, containerd, NVIDIA Container Toolkit, CUDA, cuDNN, NCCL, Fabric Manager, ROCm, RCCL, DKMS, GPUDirect RDMA, GPUDirect Storage, MOFED, DOCA, PyTorch, JAX, DCGM exporter, Prometheus, Grafana, PXE, MaaS, Packer, Foreman, Terraform, Lustre, GPFS, Weka, vLLM, Triton Inference Server, TensorRT-LLM, Nsight Systems, Nsight Compute, rocprof, perf, eBPF, EMR
3mo
Save
Mark Applied
Hide
Application Engineer
San Jose, California, United States
OnsiteFull Time
Advantest
AdvantestTokyo Stock Exchange: 6857: Manufacturer of automated test systems for the semiconductor industry.
5+ YOEBachelor’s or Master’s in Electrical Engineering, Physics, CS; 5+ years in semiconductor testing; Java; Linux; test program development; V93000 SmarTest experience; Python/Perl a plus.
Java, Linux, Perl, Python, C++, SmarTest, V93000, SoC, HPC