125 ml infrastructure engineer jobs at 85 companies in Cotati, CA

2mo
Save
Mark Applied
Hide
Founding ML infrastructure Engineer
San Francisco or United States
$200k-$350k/yr RemoteFull Time
uRun
uRun: Infrastructure cloud for interactive, stateful AI inference.
Experience designing and operating large-scale distributed infrastructure; Kubernetes/Slurm; multi-cloud GPU; reliability and scheduling; startup mindset.
Kubernetes, Slurm, Scheduling, TensorRT-LLM, NCCL, InfiniBand, RoCE, CuTe, Triton, TileLang
1mo
Save
Mark Applied
Hide
HPC/ML Infrastructure Engineer
San Francisco or Tokyo
OnsiteFull Time
Spellbrush
Spellbrush: Develops anime-themed video games using proprietary generative AI technology.
Experienced HPC/ML infrastructure engineer with Linux sysadmin skills, cluster bring-up and operations experience, familiarity with SLURM and parallel filesystems, networking and datacenter hardware handling.
SLURM, Slinky, K8s, Warewulf, MAAS, Ansible, WEKA, VAST, Ceph, Tailscale, Grafana, Prometheus, LDAP, dmesg, HGX, VLAN
6d
Save
Mark Applied
Hide
Senior Software Engineer, ML Infrastructure
Cupertino or San Francisco or Houston
$209k-$235k/yr OnsiteFull Time
Gridmatic
Gridmatic: AI-powered platform for optimizing energy trading and battery storage.
Significant experience building and operating production cloud infrastructure (GCP/AWS/Azure), Kubernetes (GKE), Terraform, workflow orchestration, Python and systems language experience; strong distributed systems and cloud networking skills.
GCP, AWS, Azure, Kubernetes, GKE, Terraform, Flyte, Temporal, Airflow, Python, Go, C++, Java, Rust, Grafana, Google Cloud Monitoring
3w
Save
Mark Applied
Hide
ML Infrastructure Engineer
San Francisco, California, United States
$180k-$230k/yr OnsiteFull Time
Echo Neurotechnologies
Echo Neurotechnologies: Developing brain-computer interface technologies to improve patient autonomy.
5+ YOEBachelor's in CS/EE or related,5+ years software or systems ML experience,proficient Python and PyTorch,distributed-training and large-scale data pipeline experience,excellent communication.
Python, PyTorch, FSDP, DeepSpeed, Megatron-LM, Ray, C++, Go, CUDA, Rust, Java, Kubernetes, Docker
4d
Save
Mark Applied
Hide
ML Engineer, Foundation Model Infrastructure
Mountain View or San Francisco or Kirkland or New York City
$175k-$215k/yr HybridFull Time
Waymo
Waymo: Autonomous driving technology for ride-hailing and logistics.
Master's degree or equivalent practical experience; Python and C++ proficiency; modern deep learning framework familiarity; experience with large-scale data pipelines or ML infrastructure.
Python, C++, Flume, JAX, PyTorch, TensorFlow, Spark, Borg, Kubeflow
3mo
Save
Mark Applied
Hide
Staff Software Engineer, ML Infrastructure
San Francisco, California, United States
$220k-$260k/yr HybridFull Time
Voxel
Voxel: AI platform for real-time industrial workplace safety monitoring.
7+ YOE7+ years software systems, ML infra experience, Python, PyTorch, ML tooling, strong communication.
Python, PyTorch, AWS, TensorRT, ONNX, Weights & Biases, MLflow, ClearML
2w
Save
Mark Applied
Hide
ML & Cloud Infrastructure Engineer Intern
South San Francisco, California, United States
OnsiteInternship
Gritt Robotics
Gritt Robotics: Developing autonomous robots to automate outdoor construction tasks.
Pursuing BS/MS/PhD in CS or related field; strong Python; familiarity with AWS/GCP, containers, CI/CD; experience building infrastructure or data systems; authorized to intern in the U.S.
Python, AWS, GCP, CI/CD, Kubernetes, Ray, Spark, Terraform
2mo
Save
Mark Applied
Hide
ML ENGINEER (GENERAL)
San Francisco, California, United States
OnsiteFull Time
MakerMaker
MakerMaker: Small San Francis-based team building autonomous ML systems
6+ YOESenior ML engineer with 6+ years building production-grade ML systems; strong Python; distributed systems experience; familiar with Ray, Kubernetes, and experimentation infrastructure.
Python, PyTorch, JAX, Ray, Slurm, Kubernetes
4d
Save
Mark Applied
Hide
AI Infrastructure Engineer
San Francisco, California, United States
$150k-$220k/yr OnsiteFull Time
Sciforium
Sciforium: Building multimodal AI models and high-performance model serving infrastructure.
5+ YOE5+ years in systems or infrastructure engineering with GPU, HPC, or ML infrastructure experience; technical bachelor's or master's degree; Linux, Kubernetes, schedulers, configuration management, Python, Bash, containers, GPUs, and RDMA expertise.
Ansible, SaltStack, Git, Python, Bash, Kubernetes, NVIDIA GPU Operator, Slurm, Run:AI, enroot, pyxis, Docker, containerd, NVIDIA Container Toolkit, CUDA, cuDNN, NCCL, Fabric Manager, ROCm, RCCL, DKMS, GPUDirect RDMA, GPUDirect Storage, MOFED, DOCA, PyTorch, JAX, DCGM exporter, Prometheus, Grafana, PXE, MaaS, Packer, Foreman, Terraform, Lustre, GPFS, Weka, vLLM, Triton Inference Server, TensorRT-LLM, Nsight Systems, Nsight Compute, rocprof, perf, eBPF, EMR
6d
Save
Mark Applied
Hide
Staff Research Engineer, Scientific Computing and ML/Physics Infrastructure
Cambridge or London or San Francisco
$224k-$294k/yr OnsiteFull Time
Lila Sciences
Lila Sciences: Develops an AI platform for autonomous scientific research and discovery.
Strong software engineering in Python, experience with ML/scientific computing, distributed systems, GPU performance, PyTorch/JAX/CUDA, Linux and containers, orchestration systems, and working with research teams.
Python, PyTorch, JAX, CUDA, Linux, Docker, Kubernetes, Slurm, Ray, Flyte, Argo, CI
2mo
Save
Mark Applied
Hide
Senior AI ML Engineer - Remote
San Francisco or Minneapolis or Washington, D.C. or United States
$120k-$215k/yr RemoteFull Time
UnitedHealth Group
UnitedHealth GroupNYSE: UNH: Provides health insurance and technology-enabled health care services.
4+ YOE2+ MgmtBachelor's degree or 4+ years equivalent, 4+ years Python, 4+ years cloud infrastructure (AWS/Azure/GCP), 4+ years AI/ML infrastructure experience, 2+ years team lead, 1+ year LLM experience.
Python, AWS, Azure, GCP, Large Language Models (LLMs), GitHub, GitHub Actions, Docker, Terraform, CI/CD
2mo
Save
Mark Applied
Hide
Infrastructure Engineer
New York or San Francisco or United States
$165k-$200k/yr HybridFull Time
Roboflow
Roboflow: Platform for building and deploying custom computer vision models.
Kubernetes production experience; IaC (Terraform/Helm); cloud (AWS/GCP); Python/Node.js; CI/CD (GitHub Actions/Spacelift); security and ML/AI infrastructure familiarity.
Kubernetes, Terraform, Helm, Python, Node.js, GitHub Actions, Spacelift, AWS, GCP, PyTorch, TensorFlow, Bash
2mo
Save
Mark Applied
Hide
Software Engineer, Monetization ML Infrastructure
San Francisco, California, United States
$293k-$441k/yr HybridFull Time
OpenAI
OpenAI: Develops artificial intelligence models and generative AI software services.
7+ YOE7+ years of professional software engineering, experience with large-scale distributed systems or ML infrastructure, built ML workflows and data pipelines, low-latency, reliable systems, observability, and cross-functional collaboration.
1w
Save
Mark Applied
Hide
Software Engineer - ML Infrastructure
San Francisco, California, United States
OnsiteFull Time
Epsilon Health
Epsilon Health: Provides AI-powered radiology diagnostic services and medical imaging software.
5+ YOE5+ years building production ML infrastructure and data pipelines, strong Python and PyTorch/JAX skills, distributed training and cloud experience, familiarity with data pipeline technologies and containerization.
Python, PyTorch, JAX, FSDP, DeepSpeed, Megatron, Spark, Airflow, BigQuery, Snowflake, Databricks, Chalk, AWS, GCP, Docker, Kubernetes, vLLM, SGLang, TensorRT, Triton, DICOM, protobufs
1mo
Save
Mark Applied
Hide
Machine Learning Infrastructure Engineer
San Francisco or New York or Los Angeles or Seattle
$200k-$345k/yr HybridFull Time
Whatnot
Whatnot: Social marketplace for buying and selling via live streams
4+ YOE4+ years building ML systems,3+ years software engineering,1+ year Python,experience with databases,monitoring,cloud services and production ML deployments.
Python, PostgreSQL, DynamoDB, Elasticsearch, Redis, DataDog, Grafana, AWS Sagemaker, Lambda, Kinesis, S3, EC2, EKS, ECS, Apache Kafka, Flink
3mo
Save
Mark Applied
Hide
SWE - Backend Infrastructure Engineer
San Francisco or Bellevue or New York
$175k-$280k/yr OnsiteFull Time
Sesame
Sesame: Designing wearable computers with lifelike voice-driven AI agents.
3+ YOEStrong systems thinker with reliability engineering experience; 3+ years in infrastructure, platform, or ML systems; Kubernetes production experience; strong communication.
Kubernetes, Terraform, CloudFormation, Pulumi, TorchServe, Seldon, KServe, Ray Serve, PyTorch, APIs, Database design
3w
Save
Mark Applied
Hide
Member of Technical Staff - Machine Learning Infrastructure Engineer
San Francisco or Toronto or Seattle
$180k-$300k/yr OnsiteFull Time
Preference Model
Preference Model: Building reinforcement learning environments to train frontier AI models.
Experienced software engineer with production ML/data infrastructure skills, proficiency with PyTorch or JAX, distributed systems, AWS/GCP, Kubernetes, data pipelines, and familiarity with transformers and inference libraries like vLLM.
PyTorch, JAX, AWS, GCP, Kubernetes, transformers, vLLM, SGLang
1w
Save
Mark Applied
Hide
Senior AI / ML Engineer
San Francisco or Melbourne
OnsiteFull Time
Artificial Analysis
Artificial Analysis: Independent AI benchmarking and performance analysis platform.
3+ YOE3+ years software engineering experience, Python and pandas proficiency, OpenAI API experience, cloud infrastructure familiarity, data visualization and communication skills; Bachelor’s or Master’s preferred.
Python, pandas, OpenAI, PyTorch
1mo
Save
Mark Applied
Hide
Staff Machine Learning Infrastructure Engineer
San Francisco, California, United States
$224k-$280k/yr OnsiteFull Time
Atoms
Atoms: Building robotics and software to automate physical world industries.
8+ YOE8+ years software engineering experience; backend systems programming in Go/Python/Java; Kubernetes, distributed ML compute (e.g., Ray), MLflow, and high-throughput data pipeline experience.
Go, Python, Java, Rust, Kubernetes, Ray, MLflow
1mo
Save
Mark Applied
Hide
Staff Machine Learning Infrastructure Engineer
San Francisco, California, United States
$224k-$280k/yr OnsiteFull Time
Atoms
Atoms: Building specialized industrial robots and physical AI systems.
8+ YOE8+ years engineering experience; strong backend systems skills (Go, Python, Java; Rust a plus); Kubernetes, distributed ML frameworks (e.g., Ray), MLflow, and high-throughput data pipeline experience.
Go, Python, Java, Rust, Kubernetes, Ray, MLflow