146 ml infrastructure engineer jobs at 95 companies in Fairfield, CA

2mo
Save
Mark Applied
Hide
Founding ML infrastructure Engineer
San Francisco or United States
$200k-$350k/yr RemoteFull Time
uRun
uRun: Infrastructure cloud for interactive, stateful AI inference.
Experience designing and operating large-scale distributed infrastructure; Kubernetes/Slurm; multi-cloud GPU; reliability and scheduling; startup mindset.
Kubernetes, Slurm, Scheduling, TensorRT-LLM, NCCL, InfiniBand, RoCE, CuTe, Triton, TileLang
1mo
Save
Mark Applied
Hide
HPC/ML Infrastructure Engineer
San Francisco or Tokyo
OnsiteFull Time
Spellbrush
Spellbrush: Develops anime-themed video games using proprietary generative AI technology.
Experienced HPC/ML infrastructure engineer with Linux sysadmin skills, cluster bring-up and operations experience, familiarity with SLURM and parallel filesystems, networking and datacenter hardware handling.
SLURM, Slinky, K8s, Warewulf, MAAS, Ansible, WEKA, VAST, Ceph, Tailscale, Grafana, Prometheus, LDAP, dmesg, HGX, VLAN
2d
Save
Mark Applied
Hide
Senior Software Engineer, ML Infrastructure
Cupertino or San Francisco or Houston
$209k-$235k/yr OnsiteFull Time
Gridmatic
Gridmatic: AI-powered platform for optimizing energy trading and battery storage.
Significant experience building and operating production cloud infrastructure (GCP/AWS/Azure), Kubernetes (GKE), Terraform, workflow orchestration, Python and systems language experience; strong distributed systems and cloud networking skills.
GCP, AWS, Azure, Kubernetes, GKE, Terraform, Flyte, Temporal, Airflow, Python, Go, C++, Java, Rust, Grafana, Google Cloud Monitoring
2w
Save
Mark Applied
Hide
ML Infrastructure Engineer
San Francisco, California, United States
$180k-$230k/yr OnsiteFull Time
Echo Neurotechnologies
Echo Neurotechnologies: Developing brain-computer interface technologies to improve patient autonomy.
5+ YOEBachelor's in CS/EE or related,5+ years software or systems ML experience,proficient Python and PyTorch,distributed-training and large-scale data pipeline experience,excellent communication.
Python, PyTorch, FSDP, DeepSpeed, Megatron-LM, Ray, C++, Go, CUDA, Rust, Java, Kubernetes, Docker
3mo
Save
Mark Applied
Hide
Staff Software Engineer, ML Infrastructure
San Francisco, California, United States
$220k-$260k/yr HybridFull Time
Voxel
Voxel: AI platform for real-time industrial workplace safety monitoring.
7+ YOE7+ years software systems, ML infra experience, Python, PyTorch, ML tooling, strong communication.
Python, PyTorch, AWS, TensorRT, ONNX, Weights & Biases, MLflow, ClearML
2w
Save
Mark Applied
Hide
ML & Cloud Infrastructure Engineer Intern
South San Francisco, California, United States
OnsiteInternship
Gritt Robotics
Gritt Robotics: Developing autonomous robots to automate outdoor construction tasks.
Pursuing BS/MS/PhD in CS or related field; strong Python; familiarity with AWS/GCP, containers, CI/CD; experience building infrastructure or data systems; authorized to intern in the U.S.
Python, AWS, GCP, CI/CD, Kubernetes, Ray, Spark, Terraform
1mo
Save
Mark Applied
Hide
Senior / Principal Infrastructure Engineer - ML Platform
San Mateo, California, United States
$279k-$345k/yr HybridFull Time
Roblox
RobloxNYSE: RBLX: Platform for creating and playing user-generated 3D digital experiences.
6+ YOE6+ years experience building scalable infrastructure; deep Kubernetes and Terraform experience; familiarity with AWS/GCP, Docker, CI/CD; bachelor's degree or equivalent practical experience.
Kubernetes, Terraform, AWS, GCP, Docker, CI/CD
2mo
Save
Mark Applied
Hide
ML ENGINEER (GENERAL)
San Francisco, California, United States
OnsiteFull Time
MakerMaker
MakerMaker: Small San Francis-based team building autonomous ML systems
6+ YOESenior ML engineer with 6+ years building production-grade ML systems; strong Python; distributed systems experience; familiar with Ray, Kubernetes, and experimentation infrastructure.
Python, PyTorch, JAX, Ray, Slurm, Kubernetes
2h
Save
Mark Applied
Hide
AI Infrastructure Engineer
San Francisco, California, United States
$150k-$220k/yr OnsiteFull Time
Sciforium
Sciforium: Building multimodal AI models and high-performance model serving infrastructure.
5+ YOE5+ years in systems or infrastructure engineering with GPU, HPC, or ML infrastructure experience; technical bachelor's or master's degree; Linux, Kubernetes, schedulers, configuration management, Python, Bash, containers, GPUs, and RDMA expertise.
Ansible, SaltStack, Git, Python, Bash, Kubernetes, NVIDIA GPU Operator, Slurm, Run:AI, enroot, pyxis, Docker, containerd, NVIDIA Container Toolkit, CUDA, cuDNN, NCCL, Fabric Manager, ROCm, RCCL, DKMS, GPUDirect RDMA, GPUDirect Storage, MOFED, DOCA, PyTorch, JAX, DCGM exporter, Prometheus, Grafana, PXE, MaaS, Packer, Foreman, Terraform, Lustre, GPFS, Weka, vLLM, Triton Inference Server, TensorRT-LLM, Nsight Systems, Nsight Compute, rocprof, perf, eBPF, EMR
2d
Save
Mark Applied
Hide
Staff Research Engineer, Scientific Computing and ML/Physics Infrastructure
Cambridge or London or San Francisco
$224k-$294k/yr OnsiteFull Time
Lila Sciences
Lila Sciences: Develops an AI platform for autonomous scientific research and discovery.
Strong software engineering in Python, experience with ML/scientific computing, distributed systems, GPU performance, PyTorch/JAX/CUDA, Linux and containers, orchestration systems, and working with research teams.
Python, PyTorch, JAX, CUDA, Linux, Docker, Kubernetes, Slurm, Ray, Flyte, Argo, CI
2mo
Save
Mark Applied
Hide
Senior AI ML Engineer - Remote
San Francisco or Minneapolis or Washington, D.C. or United States
$120k-$215k/yr RemoteFull Time
UnitedHealth Group
UnitedHealth GroupNYSE: UNH: Provides health insurance and technology-enabled health care services.
4+ YOE2+ MgmtBachelor's degree or 4+ years equivalent, 4+ years Python, 4+ years cloud infrastructure (AWS/Azure/GCP), 4+ years AI/ML infrastructure experience, 2+ years team lead, 1+ year LLM experience.
Python, AWS, Azure, GCP, Large Language Models (LLMs), GitHub, GitHub Actions, Docker, Terraform, CI/CD
2mo
Save
Mark Applied
Hide
Infrastructure Engineer
New York or San Francisco or United States
$165k-$200k/yr HybridFull Time
Roboflow
Roboflow: Platform for building and deploying custom computer vision models.
Kubernetes production experience; IaC (Terraform/Helm); cloud (AWS/GCP); Python/Node.js; CI/CD (GitHub Actions/Spacelift); security and ML/AI infrastructure familiarity.
Kubernetes, Terraform, Helm, Python, Node.js, GitHub Actions, Spacelift, AWS, GCP, PyTorch, TensorFlow, Bash
2mo
Save
Mark Applied
Hide
Software Engineer, Monetization ML Infrastructure
San Francisco, California, United States
$293k-$441k/yr HybridFull Time
OpenAI
OpenAI: Develops artificial intelligence models and generative AI software services.
7+ YOE7+ years of professional software engineering, experience with large-scale distributed systems or ML infrastructure, built ML workflows and data pipelines, low-latency, reliable systems, observability, and cross-functional collaboration.
6d
Save
Mark Applied
Hide
Software Engineer - ML Infrastructure
San Francisco, California, United States
OnsiteFull Time
Epsilon Health
Epsilon Health: Provides AI-powered radiology diagnostic services and medical imaging software.
5+ YOE5+ years building production ML infrastructure and data pipelines, strong Python and PyTorch/JAX skills, distributed training and cloud experience, familiarity with data pipeline technologies and containerization.
Python, PyTorch, JAX, FSDP, DeepSpeed, Megatron, Spark, Airflow, BigQuery, Snowflake, Databricks, Chalk, AWS, GCP, Docker, Kubernetes, vLLM, SGLang, TensorRT, Triton, DICOM, protobufs
2mo
Save
Mark Applied
Hide
Senior Machine Learning Infrastructure Engineer, Simulation
Mountain View or San Francisco
$213k-$263k/yr OnsiteFull Time
Waymo
Waymo: Autonomous driving technology for ride-hailing and logistics.
5+ YOE5+ years of professional software engineering, with 3+ years in ML infrastructure; BS in CS/Robotics or related field.
DeepSpeed, PyTorch, TensorFlow, ML infrastructure
4w
Save
Mark Applied
Hide
Machine Learning Infrastructure Engineer
San Francisco or New York or Los Angeles or Seattle
$200k-$345k/yr HybridFull Time
Whatnot
Whatnot: Social marketplace for buying and selling via live streams
4+ YOE4+ years building ML systems,3+ years software engineering,1+ year Python,experience with databases,monitoring,cloud services and production ML deployments.
Python, PostgreSQL, DynamoDB, Elasticsearch, Redis, DataDog, Grafana, AWS Sagemaker, Lambda, Kinesis, S3, EC2, EKS, ECS, Apache Kafka, Flink
3mo
Save
Mark Applied
Hide
SWE - Backend Infrastructure Engineer
San Francisco or Bellevue or New York
$175k-$280k/yr OnsiteFull Time
Sesame
Sesame: Designing wearable computers with lifelike voice-driven AI agents.
3+ YOEStrong systems thinker with reliability engineering experience; 3+ years in infrastructure, platform, or ML systems; Kubernetes production experience; strong communication.
Kubernetes, Terraform, CloudFormation, Pulumi, TorchServe, Seldon, KServe, Ray Serve, PyTorch, APIs, Database design
2mo
Save
Mark Applied
Hide
Senior AI/ML Platform Engineer
San Mateo, California, United States
$148k-$247k/yr HybridFull Time
Guidewire
GuidewireNYSE: GWRE: Provides a software platform for property and casualty insurers.
10+ YOE10+ years software engineering; 5+ years ML platforms/infrastructure; distributed systems; Python/Go/Java; Docker/Kubernetes; MLOps tools; cloud experience; knowledge of ML models.
Python, Go, Java, Docker, Kubernetes, MLflow, Kubeflow, SageMaker, Vertex AI, Databricks, AWS, GCP, Azure
2w
Save
Mark Applied
Hide
Member of Technical Staff - Machine Learning Infrastructure Engineer
San Francisco or Toronto or Seattle
$180k-$300k/yr OnsiteFull Time
Preference Model
Preference Model: Building reinforcement learning environments to train frontier AI models.
Experienced software engineer with production ML/data infrastructure skills, proficiency with PyTorch or JAX, distributed systems, AWS/GCP, Kubernetes, data pipelines, and familiarity with transformers and inference libraries like vLLM.
PyTorch, JAX, AWS, GCP, Kubernetes, transformers, vLLM, SGLang
1w
Save
Mark Applied
Hide
Senior AI / ML Engineer
San Francisco or Melbourne
OnsiteFull Time
Artificial Analysis
Artificial Analysis: Independent AI benchmarking and performance analysis platform.
3+ YOE3+ years software engineering experience, Python and pandas proficiency, OpenAI API experience, cloud infrastructure familiarity, data visualization and communication skills; Bachelor’s or Master’s preferred.
Python, pandas, OpenAI, PyTorch