361 ml infrastructure engineer jobs at 141 companies in Dublin, CA
1w
Save
Mark Applied
Hide
1w
ML Infrastructure Engineer
Palo Alto, California, United States
$180k-$440k/yrOnsiteFull Time
xAI: Develops advanced artificial intelligence systems to understand the universe.
2+ YOE2+ years building large-scale production systems or ML infrastructure; degree in CS or related field or equivalent experience; strong Python and compiled-language skills; experience with GPU and distributed systems.
General MotorsNYSE: GM: Manufactures and sells automobiles and automotive parts globally.
5+ YOE5+ years building large-scale distributed or ML systems; strong APIs and cloud infrastructure experience; expertise in ML lifecycle and MLOps; coding in Python or C++; BS/MS/PhD in CS/Math or equivalent experience.
Spellbrush: Develops anime-themed video games using proprietary generative AI technology.
Experienced HPC/ML infrastructure engineer with Linux sysadmin skills, cluster bring-up and operations experience, familiarity with SLURM and parallel filesystems, networking and datacenter hardware handling.
Nuro: Builds autonomous driving software and electric delivery robots.
3+ YOE3+ years in ML infrastructure/backend platform or distributed systems. Experience with Terraform/Pulumi/Crossplane, Kubernetes/Ray/Slurm/Volcano schedulers, Apache Spark/Beam, feature stores (Feast/Hopsworks/Redis), and systems design for HPC.
Sr./Staff ML Infrastructure Engineer, Compute (TPU Scheduling) - Foundation Model
Cupertino, California, United States
OnsiteFull Time
AppleNASDAQ: AAPL: Designs and sells consumer electronics, software, and online services.
Experience building schedulers, resource managers, or orchestration systems for distributed workloads; experience with TPU/GPU accelerator infrastructure, distributed ML training/inference, and frameworks such as JAX, PyTorch, TensorFlow, Ray, Pathways; MS/PhD preferred.
Echo Neurotechnologies: Developing brain-computer interface technologies to improve patient autonomy.
5+ YOEBachelor's in CS/EE or related,5+ years software or systems ML experience,proficient Python and PyTorch,distributed-training and large-scale data pipeline experience,excellent communication.
Gritt Robotics: Developing autonomous robots to automate outdoor construction tasks.
4+ YOE4+ years deploying high-performance ML pipelines; proficient in Python; comfortable with C++/Go; PyTorch; cloud platforms (AWS, GCP, Azure); Docker/Kubernetes/Airflow; Parquet/HDF5/TFRecord; US work authorization.
MaxInsights: Provides robot data collection for physical AI development.
Experience building production ML training and deployment systems, strong software engineering and infra fundamentals, PyTorch experience, HPC/GPU knowledge, and strong communication and product sense.
Docker: Provides a platform for building, sharing, and running containerized applications.
5+ YOE5+ years applied ML/AI experience, 4+ years software engineering, experience with LLM-based systems, model lifecycle and ML infrastructure, bachelor's in CS/Engineering or equivalent, strong communication and mentoring skills.
Senior / Principal Infrastructure Engineer - ML Platform
San Mateo, California, United States
$279k-$345k/yrHybridFull Time
RobloxNYSE: RBLX: Platform for creating and playing user-generated 3D digital experiences.
6+ YOE6+ years experience building scalable infrastructure; deep Kubernetes and Terraform experience; familiarity with AWS/GCP, Docker, CI/CD; bachelor's degree or equivalent practical experience.
Skydio: Develops autonomous AI drones for defense and industrial inspection.
Experience with data engineering, cloud ML platforms, ML Ops, large-scale data pipelines, model training/deployment, and security/compliance in ML infrastructure; strong collaboration.
Cloud platforms, Containerization, ML Ops, Databases, Data pipelines
MakerMaker: Small San Francis-based team building autonomous ML systems
6+ YOESenior ML engineer with 6+ years building production-grade ML systems; strong Python; distributed systems experience; familiar with Ray, Kubernetes, and experimentation infrastructure.
San Francisco or Minneapolis or Washington, D.C. or United States
$120k-$215k/yrRemoteFull Time
UnitedHealth GroupNYSE: UNH: Provides health insurance and technology-enabled health care services.
4+ YOE2+ MgmtBachelor's degree or 4+ years equivalent, 4+ years Python, 4+ years cloud infrastructure (AWS/Azure/GCP), 4+ years AI/ML infrastructure experience, 2+ years team lead, 1+ year LLM experience.
Python, AWS, Azure, GCP, Large Language Models (LLMs), GitHub, GitHub Actions, Docker, Terraform, CI/CD
OpenAI: Develops artificial intelligence models and generative AI software services.
7+ YOE7+ years of professional software engineering, experience with large-scale distributed systems or ML infrastructure, built ML workflows and data pipelines, low-latency, reliable systems, observability, and cross-functional collaboration.
Epsilon Health: Provides AI-powered radiology diagnostic services and medical imaging software.
5+ YOE5+ years building production ML infrastructure and data pipelines, strong Python and PyTorch/JAX skills, distributed training and cloud experience, familiarity with data pipeline technologies and containerization.
ML Systems Research Engineer, RL / Inference / Agent Systems
Santa Clara, California, United States
HybridFull Time
AMDNASDAQ: AMD: Designs and manufactures computer processors and graphics technology.
Experienced ML systems engineer with strong Python and ML framework skills, experience in RL/inference systems, distributed experimentation, and GPU/infrastructure workflows; advanced degree preferred.
Python, PyTorch, JAX, TensorFlow, Kubernetes, Ray, Slurm, ROCm, HIP, CUDA