876 ml infrastructure engineer jobs at 428 companies in United States

1mo
Save
Mark Applied
Hide
ML Infrastructure Engineer
Palo Alto, California, United States
$180k-$440k/yr OnsiteFull Time
xAI
xAI: Artificial intelligence research and development.
2+ YOE2+ years building large-scale production systems or ML infrastructure; degree in CS or related field or equivalent experience; strong Python and compiled-language skills; experience with GPU and distributed systems.
Python, C++, Rust, JAX, PyTorch, NVIDIA drivers, CUDA, Linux, Slurm, Puppet, Ansible
2mo
Save
Mark Applied
Hide
Staff ML Infrastructure Engineer
Sunnyvale, California, United States
$189k-$291k/yr HybridFull Time
General Motors
General MotorsNYSE: GM: Designing, building, and selling the world's best vehicles.
5+ YOE5+ years building large-scale distributed or ML systems; strong APIs and cloud infrastructure experience; expertise in ML lifecycle and MLOps; coding in Python or C++; BS/MS/PhD in CS/Math or equivalent experience.
Python, C++, PyTorch, TensorFlow, Bazel, Buck, Blaze, CMake, Docker, Kubernetes
5d
Save
Mark Applied
Hide
ML Infrastructure Engineer
United States
$100k-$150k/yr RemoteFull Time
Bright Vision Technologies
Bright Vision Technologies: AI-powered enterprise automation and software development firm.
6+ YOEBachelor’s or master’s degree in computer science or related field; 6+ years in infrastructure, platform, or HPC engineering; GPU cluster experience; Python and Go or C++; distributed training, cloud, Kubernetes, Linux, networking, and storage expertise.
PyTorch, JAX, DeepSpeed, FSDP, Megatron-LM, Ray Train, RDMA, InfiniBand, NCCL, Python, Go, C++, Kubernetes, Slurm, Ray, Linux, CI/CD
6d
Save
Mark Applied
Hide
ML Infrastructure Engineer
San Mateo, California, United States
OnsiteFull Time
Clera
Clera: AI-powered talent agent matching candidates to startup roles.
5+ YOE5+ years building production ML inference or model-serving systems. Requires scalable distributed systems, Docker, Kubernetes, cloud experience, observability tooling, and proficiency in Python, Go, Rust, C++, or Java.
TensorFlow Serving, TorchServe, Triton, KServe, Docker, Kubernetes, Prometheus, Grafana, ELK, AWS, GCP, Azure, Python, Go, Rust, C++, Java, Neo4j, Amazon Neptune
2mo
Save
Mark Applied
Hide
HPC/ML Infrastructure Engineer
San Francisco or Tokyo
OnsiteFull Time
Spellbrush
Spellbrush: Generative AI and anime game studio making anime illustrations and video games for artists and players.
Experienced HPC/ML infrastructure engineer with Linux sysadmin skills, cluster bring-up and operations experience, familiarity with SLURM and parallel filesystems, networking and datacenter hardware handling.
SLURM, Slinky, K8s, Warewulf, MAAS, Ansible, WEKA, VAST, Ceph, Tailscale, Grafana, Prometheus, LDAP, dmesg, HGX, VLAN
4d
Save
Mark Applied
Hide
Senior ML Infra Engineer
New York City or San Francisco
$200k-$275k/yr HybridFull Time
General Legal
General Legal: AI-native law firm serving growing companies with flat-fee commercial, corporate, and employment legal services.
5+ YOERequires 5+ years in software, machine learning, or infrastructure engineering; Python proficiency; production ML infrastructure, cloud, distributed systems, containers, independent system design, and a Computer Science degree or equivalent.
Python
2mo
Save
Mark Applied
Hide
ML Infrastructure Engineer
California, United States
OnsiteFull Time
Maven Robotics
Maven Robotics: Private robotics building general-purpose AI robots for manufacturing and logistics organizations.
Significant experience building and operating production backend, distributed, or compute infrastructure; strong programming in Python, Go, Rust, or C++; experience with Kubernetes, Ray, ZenML, storage, observability, IaC, and GPU compute orchestration.
Python, Go, Rust, C++, Kubernetes, Ray, ZenML
2mo
Save
Mark Applied
Hide
Software Engineer, ML Infrastructure
Mountain View, California, United States
$160k-$241k/yr OnsiteFull Time
Nuro
Nuro: Private U.S. autonomous-driving technology developing AI systems for automakers and mobility providers.
3+ YOE3+ years in ML infrastructure/backend platform or distributed systems. Experience with Terraform/Pulumi/Crossplane, Kubernetes/Ray/Slurm/Volcano schedulers, Apache Spark/Beam, feature stores (Feast/Hopsworks/Redis), and systems design for HPC.
Terraform, Pulumi, Crossplane, Kubernetes, KubeRay, Ray, Slurm, Volcano, Apache Spark, Apache Beam, Feast, Hopsworks, Redis, Lustre, Ceph, NVMe, AWS, GCP, Azure, Kubeflow, CNCF
4w
Save
Mark Applied
Hide
Senior Software Engineer, ML Infrastructure
Cupertino or San Francisco or Houston
$209k-$235k/yr OnsiteFull Time
Gridmatic
Gridmatic: AI-powered energy supplying and optimizing clean electricity for commercial and industrial customers.
Significant experience building and operating production cloud infrastructure (GCP/AWS/Azure), Kubernetes (GKE), Terraform, workflow orchestration, Python and systems language experience; strong distributed systems and cloud networking skills.
GCP, AWS, Azure, Kubernetes, GKE, Terraform, Flyte, Temporal, Airflow, Python, Go, C++, Java, Rust, Grafana, Google Cloud Monitoring
8h
Save
Mark Applied
Hide
AIML - Staff ML Infrastructure Engineer, ML Platform & Technology - Pre-training Infrastructure
California, United States
OnsiteFull Time
Apple
AppleNASDAQ: AAPL: Designing and manufacturing consumer electronics, software, and digital services.
6+ YOE6+ years building or optimizing high-performance ML or distributed systems; programming proficiency; distributed systems and parallel computing expertise; profiling experience; bachelor's degree in computer science, engineering, or related field.
Python, TPU, JAX, XLA, PyTorch, Pallas, Triton, CUDA, GPU
1mo
Save
Mark Applied
Hide
ML Infrastructure Engineer
San Francisco, California, United States
$180k-$230k/yr OnsiteFull Time
Echo Neurotechnologies
Echo Neurotechnologies: Private San Francis neurotechnology startup developing brain-computer interfaces that restore communication for people with severe disabilities.
5+ YOEBachelor's in CS/EE or related,5+ years software or systems ML experience,proficient Python and PyTorch,distributed-training and large-scale data pipeline experience,excellent communication.
Python, PyTorch, FSDP, DeepSpeed, Megatron-LM, Ray, C++, Go, CUDA, Rust, Java, Kubernetes, Docker
1mo
Save
Mark Applied
Hide
Senior ML Ops Engineer (Machine Learning Infrastructure)
Los Angeles, California, United States
$150k-$250k/yr HybridFull Time
Parallel Systems
Parallel Systems: U.S. manufacturer of autonomous battery-electric rail vehicles serving railroads and short-haul freight markets.
5+ YOE5+ years building large-scale systems with 2+ years on ML infrastructure; BS in CS/ML/engineering; proficiency in Python and Git; experience with CI/CD, distributed training, cloud ML architectures; strong communication and system design skills.
MLflow, SageMaker, Kubeflow, Airflow, Metaflow, Python, Git, PyTorch DDP, Horovod, Ray, AWS, GCP, Azure
3w
Save
Mark Applied
Hide
ML Engineer, Foundation Model Infrastructure
Mountain View or San Francisco or Kirkland or New York City
$175k-$215k/yr HybridFull Time
Waymo
Waymo: Autonomous driving technology and robotaxi service provider.
Master's degree or equivalent practical experience; Python and C++ proficiency; modern deep learning framework familiarity; experience with large-scale data pipelines or ML infrastructure.
Python, C++, Flume, JAX, PyTorch, TensorFlow, Spark, Borg, Kubeflow
1mo
Save
Mark Applied
Hide
ML & Cloud Infrastructure Engineer Intern
South San Francisco, California, United States
OnsiteInternship
Gritt Robotics
Gritt Robotics: Private AI-powered construction robotics automating labor-intensive infrastructure work for construction crews.
Pursuing BS/MS/PhD in CS or related field; strong Python; familiarity with AWS/GCP, containers, CI/CD; experience building infrastructure or data systems; authorized to intern in the U.S.
Python, AWS, GCP, CI/CD, Kubernetes, Ray, Spark, Terraform
4w
Save
Mark Applied
Hide
Senior Data & ML Infrastructure Engineer (Xora Portfolio Company)
Singapore or United States
HybridFull Time
Xora Innovation
Xora Innovation: Singapore-based private venture capital platform of Temasek investing in and building early-stage AI and deep-tech startups.
6+ YOEBachelor's or Master's in CS or related,6+ years building production software, strong Python, experience with large-scale data systems, ML infra, containers and orchestration, workflow orchestrators, and telemetry/monitoring.
Python, Airflow, Dagster, Flyte, Temporal, Docker, Kubernetes, Prometheus, Grafana, OpenTelemetry, Great Expectations, Evidently, MLflow, Weights & Biases, Ray Serve, KServe, Kubeflow, Atompack, ASE
4h
Save
Mark Applied
Hide
ML Infra Engineer, Modeling
San Francisco, California, United States
OnsiteFull Time
Physical Intelligence
Physical Intelligence: AI robotics developing foundation models and learning algorithms for robots and physically actuated devices.
Strong software engineering fundamentals; ML training infrastructure experience; large-scale JAX or PyTorch training; distributed systems, cloud workloads, performance optimization, and cross-functional communication experience.
JAX, PyTorch, SLURM, Kubernetes, Google Cloud Platform (GCP), GKE, AWS, TPU, GPU
2mo
Save
Mark Applied
Hide
ML Engineer
Palo Alto or Seattle or Paris
$139k-$226k/yr RemoteFull Time
Docker
Docker: Privately held container application platform helping developers build, share, and run applications.
5+ YOE5+ years applied ML/AI experience, 4+ years software engineering, experience with LLM-based systems, model lifecycle and ML infrastructure, bachelor's in CS/Engineering or equivalent, strong communication and mentoring skills.
Docker Desktop, Docker Hub, Docker Scout, LLM, MCP, Agentic Platform
1mo
Save
Mark Applied
Hide
Senior / Principal Infrastructure Engineer - ML Platform
San Mateo, California, United States
$279k-$345k/yr HybridFull Time
Roblox
RobloxNYSE: RBLX: Global platform for user-created immersive digital experiences.
6+ YOE6+ years experience building scalable infrastructure; deep Kubernetes and Terraform experience; familiarity with AWS/GCP, Docker, CI/CD; bachelor's degree or equivalent practical experience.
Kubernetes, Terraform, AWS, GCP, Docker, CI/CD
2w
Save
Mark Applied
Hide
MLOps / ML Platform Engineer
United States or Canada
$170k-$200k/yr RemoteFull Time
SumerSports
SumerSports: AI-powered football intelligence serving professional and collegiate teams with scouting, roster, and performance analytics.
4+ YOERequires 4+ years in ML platform, DevOps, or infrastructure engineering; Kubernetes, CI/CD, containers, cloud, GPU clusters, Python, infrastructure as code, automation, observability, and production ML systems experience.
Kubernetes, CI/CD, AWS, GCP, Azure, Delta, Parquet, Polars, Spark, Python, REST, gRPC, LLM
3w
Save
Mark Applied
Hide
Senior Infrastructure Engineer, Research
Singapore or London or New York City or Singapore
HybridFull Time
PhysicsX
PhysicsX: Private physics-AI software helping industrial engineering and manufacturing teams design and optimize hardware.
5+ YOERequires 5+ years building and operating ML infrastructure at scale, distributed training expertise, Linux and networking fundamentals, Kubernetes and SLURM, Python, ML frameworks, and cloud GPU infrastructure experience.
NVIDIA DGX B200, NCCL, FSDP, DDP, Linux, NVLink, InfiniBand, Kubernetes, SLURM, Python, PyTorch, CoreWeave, Weights & Biases, MLflow, Prometheus, Grafana