487 ml infrastructure engineer jobs at 202 companies in California

1mo
Save
Mark Applied
Hide
ML Infrastructure Engineer
Palo Alto, California, United States
$180k-$440k/yr OnsiteFull Time
xAI
xAI: Artificial intelligence research and development.
2+ YOE2+ years building large-scale production systems or ML infrastructure; degree in CS or related field or equivalent experience; strong Python and compiled-language skills; experience with GPU and distributed systems.
Python, C++, Rust, JAX, PyTorch, NVIDIA drivers, CUDA, Linux, Slurm, Puppet, Ansible
2mo
Save
Mark Applied
Hide
Staff ML Infrastructure Engineer
Sunnyvale, California, United States
$189k-$291k/yr HybridFull Time
General Motors
General MotorsNYSE: GM: Designing, building, and selling the world's best vehicles.
5+ YOE5+ years building large-scale distributed or ML systems; strong APIs and cloud infrastructure experience; expertise in ML lifecycle and MLOps; coding in Python or C++; BS/MS/PhD in CS/Math or equivalent experience.
Python, C++, PyTorch, TensorFlow, Bazel, Buck, Blaze, CMake, Docker, Kubernetes
6d
Save
Mark Applied
Hide
ML Infrastructure Engineer
San Mateo, California, United States
OnsiteFull Time
Clera
Clera: AI-powered talent agent matching candidates to startup roles.
5+ YOE5+ years building production ML inference or model-serving systems. Requires scalable distributed systems, Docker, Kubernetes, cloud experience, observability tooling, and proficiency in Python, Go, Rust, C++, or Java.
TensorFlow Serving, TorchServe, Triton, KServe, Docker, Kubernetes, Prometheus, Grafana, ELK, AWS, GCP, Azure, Python, Go, Rust, C++, Java, Neo4j, Amazon Neptune
2mo
Save
Mark Applied
Hide
HPC/ML Infrastructure Engineer
San Francisco or Tokyo
OnsiteFull Time
Spellbrush
Spellbrush: Generative AI and anime game studio making anime illustrations and video games for artists and players.
Experienced HPC/ML infrastructure engineer with Linux sysadmin skills, cluster bring-up and operations experience, familiarity with SLURM and parallel filesystems, networking and datacenter hardware handling.
SLURM, Slinky, K8s, Warewulf, MAAS, Ansible, WEKA, VAST, Ceph, Tailscale, Grafana, Prometheus, LDAP, dmesg, HGX, VLAN
4d
Save
Mark Applied
Hide
Senior ML Infra Engineer
New York City or San Francisco
$200k-$275k/yr HybridFull Time
General Legal
General Legal: AI-native law firm serving growing companies with flat-fee commercial, corporate, and employment legal services.
5+ YOERequires 5+ years in software, machine learning, or infrastructure engineering; Python proficiency; production ML infrastructure, cloud, distributed systems, containers, independent system design, and a Computer Science degree or equivalent.
Python
2mo
Save
Mark Applied
Hide
ML Infrastructure Engineer
California, United States
OnsiteFull Time
Maven Robotics
Maven Robotics: Private robotics building general-purpose AI robots for manufacturing and logistics organizations.
Significant experience building and operating production backend, distributed, or compute infrastructure; strong programming in Python, Go, Rust, or C++; experience with Kubernetes, Ray, ZenML, storage, observability, IaC, and GPU compute orchestration.
Python, Go, Rust, C++, Kubernetes, Ray, ZenML
2mo
Save
Mark Applied
Hide
Software Engineer, ML Infrastructure
Mountain View, California, United States
$160k-$241k/yr OnsiteFull Time
Nuro
Nuro: Private U.S. autonomous-driving technology developing AI systems for automakers and mobility providers.
3+ YOE3+ years in ML infrastructure/backend platform or distributed systems. Experience with Terraform/Pulumi/Crossplane, Kubernetes/Ray/Slurm/Volcano schedulers, Apache Spark/Beam, feature stores (Feast/Hopsworks/Redis), and systems design for HPC.
Terraform, Pulumi, Crossplane, Kubernetes, KubeRay, Ray, Slurm, Volcano, Apache Spark, Apache Beam, Feast, Hopsworks, Redis, Lustre, Ceph, NVMe, AWS, GCP, Azure, Kubeflow, CNCF
4w
Save
Mark Applied
Hide
Senior Software Engineer, ML Infrastructure
Cupertino or San Francisco or Houston
$209k-$235k/yr OnsiteFull Time
Gridmatic
Gridmatic: AI-powered energy supplying and optimizing clean electricity for commercial and industrial customers.
Significant experience building and operating production cloud infrastructure (GCP/AWS/Azure), Kubernetes (GKE), Terraform, workflow orchestration, Python and systems language experience; strong distributed systems and cloud networking skills.
GCP, AWS, Azure, Kubernetes, GKE, Terraform, Flyte, Temporal, Airflow, Python, Go, C++, Java, Rust, Grafana, Google Cloud Monitoring
6h
Save
Mark Applied
Hide
AIML - Staff ML Infrastructure Engineer, ML Platform & Technology - Pre-training Infrastructure
California, United States
OnsiteFull Time
Apple
AppleNASDAQ: AAPL: Designing and manufacturing consumer electronics, software, and digital services.
6+ YOE6+ years building or optimizing high-performance ML or distributed systems; programming proficiency; distributed systems and parallel computing expertise; profiling experience; bachelor's degree in computer science, engineering, or related field.
Python, TPU, JAX, XLA, PyTorch, Pallas, Triton, CUDA, GPU
1mo
Save
Mark Applied
Hide
ML Infrastructure Engineer
San Francisco, California, United States
$180k-$230k/yr OnsiteFull Time
Echo Neurotechnologies
Echo Neurotechnologies: Private San Francis neurotechnology startup developing brain-computer interfaces that restore communication for people with severe disabilities.
5+ YOEBachelor's in CS/EE or related,5+ years software or systems ML experience,proficient Python and PyTorch,distributed-training and large-scale data pipeline experience,excellent communication.
Python, PyTorch, FSDP, DeepSpeed, Megatron-LM, Ray, C++, Go, CUDA, Rust, Java, Kubernetes, Docker
1mo
Save
Mark Applied
Hide
Senior ML Ops Engineer (Machine Learning Infrastructure)
Los Angeles, California, United States
$150k-$250k/yr HybridFull Time
Parallel Systems
Parallel Systems: U.S. manufacturer of autonomous battery-electric rail vehicles serving railroads and short-haul freight markets.
5+ YOE5+ years building large-scale systems with 2+ years on ML infrastructure; BS in CS/ML/engineering; proficiency in Python and Git; experience with CI/CD, distributed training, cloud ML architectures; strong communication and system design skills.
MLflow, SageMaker, Kubeflow, Airflow, Metaflow, Python, Git, PyTorch DDP, Horovod, Ray, AWS, GCP, Azure
3w
Save
Mark Applied
Hide
ML Engineer, Foundation Model Infrastructure
Mountain View or San Francisco or Kirkland or New York City
$175k-$215k/yr HybridFull Time
Waymo
Waymo: Autonomous driving technology and robotaxi service provider.
Master's degree or equivalent practical experience; Python and C++ proficiency; modern deep learning framework familiarity; experience with large-scale data pipelines or ML infrastructure.
Python, C++, Flume, JAX, PyTorch, TensorFlow, Spark, Borg, Kubeflow
1mo
Save
Mark Applied
Hide
ML & Cloud Infrastructure Engineer Intern
South San Francisco, California, United States
OnsiteInternship
Gritt Robotics
Gritt Robotics: Private AI-powered construction robotics automating labor-intensive infrastructure work for construction crews.
Pursuing BS/MS/PhD in CS or related field; strong Python; familiarity with AWS/GCP, containers, CI/CD; experience building infrastructure or data systems; authorized to intern in the U.S.
Python, AWS, GCP, CI/CD, Kubernetes, Ray, Spark, Terraform
2mo
Save
Mark Applied
Hide
ML Engineer
Palo Alto or Seattle or Paris
$139k-$226k/yr RemoteFull Time
Docker
Docker: Privately held container application platform helping developers build, share, and run applications.
5+ YOE5+ years applied ML/AI experience, 4+ years software engineering, experience with LLM-based systems, model lifecycle and ML infrastructure, bachelor's in CS/Engineering or equivalent, strong communication and mentoring skills.
Docker Desktop, Docker Hub, Docker Scout, LLM, MCP, Agentic Platform
1mo
Save
Mark Applied
Hide
Senior / Principal Infrastructure Engineer - ML Platform
San Mateo, California, United States
$279k-$345k/yr HybridFull Time
Roblox
RobloxNYSE: RBLX: Global platform for user-created immersive digital experiences.
6+ YOE6+ years experience building scalable infrastructure; deep Kubernetes and Terraform experience; familiarity with AWS/GCP, Docker, CI/CD; bachelor's degree or equivalent practical experience.
Kubernetes, Terraform, AWS, GCP, Docker, CI/CD
3w
Save
Mark Applied
Hide
AI Infrastructure Engineer
San Francisco, California, United States
$150k-$220k/yr OnsiteFull Time
Sciforium
Sciforium: AI infrastructure building multimodal models and high-efficiency serving software for developers and teams.
5+ YOE5+ years in systems or infrastructure engineering with GPU, HPC, or ML infrastructure experience; technical bachelor's or master's degree; Linux, Kubernetes, schedulers, configuration management, Python, Bash, containers, GPUs, and RDMA expertise.
Ansible, SaltStack, Git, Python, Bash, Kubernetes, NVIDIA GPU Operator, Slurm, Run:AI, enroot, pyxis, Docker, containerd, NVIDIA Container Toolkit, CUDA, cuDNN, NCCL, Fabric Manager, ROCm, RCCL, DKMS, GPUDirect RDMA, GPUDirect Storage, MOFED, DOCA, PyTorch, JAX, DCGM exporter, Prometheus, Grafana, PXE, MaaS, Packer, Foreman, Terraform, Lustre, GPFS, Weka, vLLM, Triton Inference Server, TensorRT-LLM, Nsight Systems, Nsight Compute, rocprof, perf, eBPF, EMR
4w
Save
Mark Applied
Hide
Senior Data & ML Infrastructure Engineer (Xora Portfolio Company)
San Diego or Singapore
HybridFull Time
Xora Innovation
Xora Innovation: Singapore-based private venture capital platform of Temasek investing in and building early-stage AI and deep-tech startups.
6+ YOERequires 6+ years building production software, strong Python, experience with large-scale data systems, ML data pipelines, MLOps, containers/orchestration, and production telemetry/monitoring.
Python, Docker, Kubernetes, Airflow, Dagster, Flyte, Temporal, Prometheus, Grafana, OpenTelemetry, Great Expectations, Evidently, MLflow, Weights & Biases, Ray Serve, KServe, Kubeflow, Atompack, ASE
4d
Save
Mark Applied
Hide
Staff ML Engineer
Pune or Bangalore or Santa Clara
HybridFull Time
Cohesity
Cohesity: AI-powered data security and management solutions provider.
10+ YOERequires 10+ years software engineering, 8+ years distributed backend systems, 3+ years LLM/RAG or ML infrastructure, production Kubernetes experience, technical leadership, and a bachelor's degree in a technical field.
Python, Java, Go, JavaScript, LLMs, Prompt Engineering, Retrieval-Augmented Generation (RAG), Kubernetes, OpenShift, EKS, GKE, Docker, Helm, Terraform, AWS, gRPC, REST APIs, SQL, NoSQL, PostgreSQL, Redis, Elasticsearch, CI/CD, DevOps, Git, Node.js, Jenkins
2w
Save
Mark Applied
Hide
Senior ML Engineer
San Francisco or Los Angeles or Denver or Austin or Chicago or New York City or Canada or Seattle or Santa Barbara or San Diego or Toronto
$152k-$228k/yr RemoteFull Time
Invoca
Invoca: Privately held SaaS platform helping marketing, commerce, and contact-center teams convert customer conversations into revenue.
5+ YOE5+ years of ML engineering experience; advanced Python, PyTorch, and deep learning; production NLP model deployment; SLM/LLM fine-tuning; inference infrastructure; production APIs; MLOps and model monitoring.
Python, PyTorch, HuggingFace Transformers, spaCy, Triton Inference Server, Baseten, Kubernetes, LoRA, QLoRA, PEFT, vLLM, TGI, SageMaker, Vertex AI, Braintrust, MLflow, RLHF
4w
Save
Mark Applied
Hide
Staff Research Engineer, Scientific Computing and ML/Physics Infrastructure
Cambridge or London or San Francisco
$224k-$294k/yr OnsiteFull Time
Lila Sciences
Lila Sciences: Builds AI and autonomous labs for scientific discovery.
Strong software engineering in Python, experience with ML/scientific computing, distributed systems, GPU performance, PyTorch/JAX/CUDA, Linux and containers, orchestration systems, and working with research teams.
Python, PyTorch, JAX, CUDA, Linux, Docker, Kubernetes, Slurm, Ray, Flyte, Argo, CI