361 ml infrastructure engineer jobs at 141 companies in Dublin, CA

1w
Save
Mark Applied
Hide
ML Infrastructure Engineer
Palo Alto, California, United States
$180k-$440k/yr OnsiteFull Time
xAI
xAI: Develops advanced artificial intelligence systems to understand the universe.
2+ YOE2+ years building large-scale production systems or ML infrastructure; degree in CS or related field or equivalent experience; strong Python and compiled-language skills; experience with GPU and distributed systems.
Python, C++, Rust, JAX, PyTorch, NVIDIA drivers, CUDA, Linux, Slurm, Puppet, Ansible
1mo
Save
Mark Applied
Hide
Staff ML Infrastructure Engineer
Sunnyvale, California, United States
$189k-$291k/yr HybridFull Time
General Motors
General MotorsNYSE: GM: Manufactures and sells automobiles and automotive parts globally.
5+ YOE5+ years building large-scale distributed or ML systems; strong APIs and cloud infrastructure experience; expertise in ML lifecycle and MLOps; coding in Python or C++; BS/MS/PhD in CS/Math or equivalent experience.
Python, C++, PyTorch, TensorFlow, Bazel, Buck, Blaze, CMake, Docker, Kubernetes
2mo
Save
Mark Applied
Hide
Founding ML infrastructure Engineer
San Francisco or United States
$200k-$350k/yr RemoteFull Time
uRun
uRun: Infrastructure cloud for interactive, stateful AI inference.
Experience designing and operating large-scale distributed infrastructure; Kubernetes/Slurm; multi-cloud GPU; reliability and scheduling; startup mindset.
Kubernetes, Slurm, Scheduling, TensorRT-LLM, NCCL, InfiniBand, RoCE, CuTe, Triton, TileLang
1mo
Save
Mark Applied
Hide
HPC/ML Infrastructure Engineer
San Francisco or Tokyo
OnsiteFull Time
Spellbrush
Spellbrush: Develops anime-themed video games using proprietary generative AI technology.
Experienced HPC/ML infrastructure engineer with Linux sysadmin skills, cluster bring-up and operations experience, familiarity with SLURM and parallel filesystems, networking and datacenter hardware handling.
SLURM, Slinky, K8s, Warewulf, MAAS, Ansible, WEKA, VAST, Ceph, Tailscale, Grafana, Prometheus, LDAP, dmesg, HGX, VLAN
1mo
Save
Mark Applied
Hide
Software Engineer, ML Infrastructure
Mountain View, California, United States
$160k-$241k/yr OnsiteFull Time
Nuro
Nuro: Builds autonomous driving software and electric delivery robots.
3+ YOE3+ years in ML infrastructure/backend platform or distributed systems. Experience with Terraform/Pulumi/Crossplane, Kubernetes/Ray/Slurm/Volcano schedulers, Apache Spark/Beam, feature stores (Feast/Hopsworks/Redis), and systems design for HPC.
Terraform, Pulumi, Crossplane, Kubernetes, KubeRay, Ray, Slurm, Volcano, Apache Spark, Apache Beam, Feast, Hopsworks, Redis, Lustre, Ceph, NVMe, AWS, GCP, Azure, Kubeflow, CNCF
1w
Save
Mark Applied
Hide
Sr./Staff ML Infrastructure Engineer, Compute (TPU Scheduling) - Foundation Model
Cupertino, California, United States
OnsiteFull Time
Apple
AppleNASDAQ: AAPL: Designs and sells consumer electronics, software, and online services.
Experience building schedulers, resource managers, or orchestration systems for distributed workloads; experience with TPU/GPU accelerator infrastructure, distributed ML training/inference, and frameworks such as JAX, PyTorch, TensorFlow, Ray, Pathways; MS/PhD preferred.
TPU, GPU, JAX, PyTorch, TensorFlow, Ray, Pathways
2w
Save
Mark Applied
Hide
ML Infrastructure Engineer
San Francisco, California, United States
$180k-$230k/yr OnsiteFull Time
Echo Neurotechnologies
Echo Neurotechnologies: Developing brain-computer interface technologies to improve patient autonomy.
5+ YOEBachelor's in CS/EE or related,5+ years software or systems ML experience,proficient Python and PyTorch,distributed-training and large-scale data pipeline experience,excellent communication.
Python, PyTorch, FSDP, DeepSpeed, Megatron-LM, Ray, C++, Go, CUDA, Rust, Java, Kubernetes, Docker
2mo
Save
Mark Applied
Hide
ML & Cloud Infrastructure Engineer
Belmont, California, United States
OnsiteFull Time
Gritt Robotics
Gritt Robotics: Developing autonomous robots to automate outdoor construction tasks.
4+ YOE4+ years deploying high-performance ML pipelines; proficient in Python; comfortable with C++/Go; PyTorch; cloud platforms (AWS, GCP, Azure); Docker/Kubernetes/Airflow; Parquet/HDF5/TFRecord; US work authorization.
Python, C++, Go, PyTorch, Parquet, HDF5, TFRecord, AWS, GCP, Azure, Docker, Kubernetes, Airflow
2w
Save
Mark Applied
Hide
Senior ML Infra Engineer
Santa Clara, California, United States
OnsiteFull Time
MaxInsights
MaxInsights: Provides robot data collection for physical AI development.
Experience building production ML training and deployment systems, strong software engineering and infra fundamentals, PyTorch experience, HPC/GPU knowledge, and strong communication and product sense.
Python, PyTorch, Docker, CI/CD, GPUs, CUDA, vLLM, TensorRT, Triton
2mo
Save
Mark Applied
Hide
Staff Software Engineer, ML Infrastructure
San Francisco, California, United States
$220k-$260k/yr HybridFull Time
Voxel
Voxel: AI platform for real-time industrial workplace safety monitoring.
7+ YOE7+ years software systems, ML infra experience, Python, PyTorch, ML tooling, strong communication.
Python, PyTorch, AWS, TensorRT, ONNX, Weights & Biases, MLflow, ClearML
1mo
Save
Mark Applied
Hide
ML Engineer
Palo Alto or Seattle or Paris
$139k-$226k/yr RemoteFull Time
Docker
Docker: Provides a platform for building, sharing, and running containerized applications.
5+ YOE5+ years applied ML/AI experience, 4+ years software engineering, experience with LLM-based systems, model lifecycle and ML infrastructure, bachelor's in CS/Engineering or equivalent, strong communication and mentoring skills.
Docker Desktop, Docker Hub, Docker Scout, LLM, MCP, Agentic Platform
3w
Save
Mark Applied
Hide
Senior / Principal Infrastructure Engineer - ML Platform
San Mateo, California, United States
$279k-$345k/yr HybridFull Time
Roblox
RobloxNYSE: RBLX: Platform for creating and playing user-generated 3D digital experiences.
6+ YOE6+ years experience building scalable infrastructure; deep Kubernetes and Terraform experience; familiarity with AWS/GCP, Docker, CI/CD; bachelor's degree or equivalent practical experience.
Kubernetes, Terraform, AWS, GCP, Docker, CI/CD
2mo
Save
Mark Applied
Hide
Autonomy Engineer - ML & DL Infrastructure
San Mateo, California, United States
$170k-$278k/yr HybridFull Time
Skydio
Skydio: Develops autonomous AI drones for defense and industrial inspection.
Experience with data engineering, cloud ML platforms, ML Ops, large-scale data pipelines, model training/deployment, and security/compliance in ML infrastructure; strong collaboration.
Cloud platforms, Containerization, ML Ops, Databases, Data pipelines
2mo
Save
Mark Applied
Hide
ML ENGINEER (GENERAL)
San Francisco, California, United States
OnsiteFull Time
MakerMaker
MakerMaker: Small San Francis-based team building autonomous ML systems
6+ YOESenior ML engineer with 6+ years building production-grade ML systems; strong Python; distributed systems experience; familiar with Ray, Kubernetes, and experimentation infrastructure.
Python, PyTorch, JAX, Ray, Slurm, Kubernetes
3w
Save
Mark Applied
Hide
Sr. Mgr., ML Infrastructure, PV Personalization and Discovery
Sunnyvale or Seattle or New York
$242k-$328k/yr OnsiteFull Time
Amazon
AmazonNASDAQ: AMZN: Global online retail and cloud computing technology provider.
10+ YOE5+ Mgmt10+ years engineering experience, 5+ years managing engineering teams, expertise in ML infrastructure, retrieval/recommendation systems, LLMs/generative AI, partnering with applied scientists, experience delivering large-scale consumer software.
AWS
2mo
Save
Mark Applied
Hide
Senior AI ML Engineer - Remote
San Francisco or Minneapolis or Washington, D.C. or United States
$120k-$215k/yr RemoteFull Time
UnitedHealth Group
UnitedHealth GroupNYSE: UNH: Provides health insurance and technology-enabled health care services.
4+ YOE2+ MgmtBachelor's degree or 4+ years equivalent, 4+ years Python, 4+ years cloud infrastructure (AWS/Azure/GCP), 4+ years AI/ML infrastructure experience, 2+ years team lead, 1+ year LLM experience.
Python, AWS, Azure, GCP, Large Language Models (LLMs), GitHub, GitHub Actions, Docker, Terraform, CI/CD
2mo
Save
Mark Applied
Hide
Infrastructure Engineer
New York or San Francisco or United States
$165k-$200k/yr HybridFull Time
Roboflow
Roboflow: Platform for building and deploying custom computer vision models.
Kubernetes production experience; IaC (Terraform/Helm); cloud (AWS/GCP); Python/Node.js; CI/CD (GitHub Actions/Spacelift); security and ML/AI infrastructure familiarity.
Kubernetes, Terraform, Helm, Python, Node.js, GitHub Actions, Spacelift, AWS, GCP, PyTorch, TensorFlow, Bash
2mo
Save
Mark Applied
Hide
Software Engineer, Monetization ML Infrastructure
San Francisco, California, United States
$293k-$441k/yr HybridFull Time
OpenAI
OpenAI: Develops artificial intelligence models and generative AI software services.
7+ YOE7+ years of professional software engineering, experience with large-scale distributed systems or ML infrastructure, built ML workflows and data pipelines, low-latency, reliable systems, observability, and cross-functional collaboration.
2d
Save
Mark Applied
Hide
Software Engineer - ML Infrastructure
San Francisco, California, United States
OnsiteFull Time
Epsilon Health
Epsilon Health: Provides AI-powered radiology diagnostic services and medical imaging software.
5+ YOE5+ years building production ML infrastructure and data pipelines, strong Python and PyTorch/JAX skills, distributed training and cloud experience, familiarity with data pipeline technologies and containerization.
Python, PyTorch, JAX, FSDP, DeepSpeed, Megatron, Spark, Airflow, BigQuery, Snowflake, Databricks, Chalk, AWS, GCP, Docker, Kubernetes, vLLM, SGLang, TensorRT, Triton, DICOM, protobufs
4d
Save
Mark Applied
Hide
ML Systems Research Engineer, RL / Inference / Agent Systems
Santa Clara, California, United States
HybridFull Time
AMD
AMDNASDAQ: AMD: Designs and manufactures computer processors and graphics technology.
Experienced ML systems engineer with strong Python and ML framework skills, experience in RL/inference systems, distributed experimentation, and GPU/infrastructure workflows; advanced degree preferred.
Python, PyTorch, JAX, TensorFlow, Kubernetes, Ray, Slurm, ROCm, HIP, CUDA