876 ml infrastructure engineer jobs at 428 companies in United States
1mo
Save
Mark Applied
Hide
1mo
ML Infrastructure Engineer
Palo Alto, California, United States
$180k-$440k/yrOnsiteFull Time
xAI: Artificial intelligence research and development.
2+ YOE2+ years building large-scale production systems or ML infrastructure; degree in CS or related field or equivalent experience; strong Python and compiled-language skills; experience with GPU and distributed systems.
General MotorsNYSE: GM: Designing, building, and selling the world's best vehicles.
5+ YOE5+ years building large-scale distributed or ML systems; strong APIs and cloud infrastructure experience; expertise in ML lifecycle and MLOps; coding in Python or C++; BS/MS/PhD in CS/Math or equivalent experience.
Bright Vision Technologies: AI-powered enterprise automation and software development firm.
6+ YOEBachelor’s or master’s degree in computer science or related field; 6+ years in infrastructure, platform, or HPC engineering; GPU cluster experience; Python and Go or C++; distributed training, cloud, Kubernetes, Linux, networking, and storage expertise.
Clera: AI-powered talent agent matching candidates to startup roles.
5+ YOE5+ years building production ML inference or model-serving systems. Requires scalable distributed systems, Docker, Kubernetes, cloud experience, observability tooling, and proficiency in Python, Go, Rust, C++, or Java.
Spellbrush: Generative AI and anime game studio making anime illustrations and video games for artists and players.
Experienced HPC/ML infrastructure engineer with Linux sysadmin skills, cluster bring-up and operations experience, familiarity with SLURM and parallel filesystems, networking and datacenter hardware handling.
General Legal: AI-native law firm serving growing companies with flat-fee commercial, corporate, and employment legal services.
5+ YOERequires 5+ years in software, machine learning, or infrastructure engineering; Python proficiency; production ML infrastructure, cloud, distributed systems, containers, independent system design, and a Computer Science degree or equivalent.
Maven Robotics: Private robotics building general-purpose AI robots for manufacturing and logistics organizations.
Significant experience building and operating production backend, distributed, or compute infrastructure; strong programming in Python, Go, Rust, or C++; experience with Kubernetes, Ray, ZenML, storage, observability, IaC, and GPU compute orchestration.
Nuro: Private U.S. autonomous-driving technology developing AI systems for automakers and mobility providers.
3+ YOE3+ years in ML infrastructure/backend platform or distributed systems. Experience with Terraform/Pulumi/Crossplane, Kubernetes/Ray/Slurm/Volcano schedulers, Apache Spark/Beam, feature stores (Feast/Hopsworks/Redis), and systems design for HPC.
Gridmatic: AI-powered energy supplying and optimizing clean electricity for commercial and industrial customers.
Significant experience building and operating production cloud infrastructure (GCP/AWS/Azure), Kubernetes (GKE), Terraform, workflow orchestration, Python and systems language experience; strong distributed systems and cloud networking skills.
AIML - Staff ML Infrastructure Engineer, ML Platform & Technology - Pre-training Infrastructure
California, United States
OnsiteFull Time
AppleNASDAQ: AAPL: Designing and manufacturing consumer electronics, software, and digital services.
6+ YOE6+ years building or optimizing high-performance ML or distributed systems; programming proficiency; distributed systems and parallel computing expertise; profiling experience; bachelor's degree in computer science, engineering, or related field.
Echo Neurotechnologies: Private San Francis neurotechnology startup developing brain-computer interfaces that restore communication for people with severe disabilities.
5+ YOEBachelor's in CS/EE or related,5+ years software or systems ML experience,proficient Python and PyTorch,distributed-training and large-scale data pipeline experience,excellent communication.
Senior ML Ops Engineer (Machine Learning Infrastructure)
Los Angeles, California, United States
$150k-$250k/yrHybridFull Time
Parallel Systems: U.S. manufacturer of autonomous battery-electric rail vehicles serving railroads and short-haul freight markets.
5+ YOE5+ years building large-scale systems with 2+ years on ML infrastructure; BS in CS/ML/engineering; proficiency in Python and Git; experience with CI/CD, distributed training, cloud ML architectures; strong communication and system design skills.
Mountain View or San Francisco or Kirkland or New York City
$175k-$215k/yrHybridFull Time
Waymo: Autonomous driving technology and robotaxi service provider.
Master's degree or equivalent practical experience; Python and C++ proficiency; modern deep learning framework familiarity; experience with large-scale data pipelines or ML infrastructure.
Gritt Robotics: Private AI-powered construction robotics automating labor-intensive infrastructure work for construction crews.
Pursuing BS/MS/PhD in CS or related field; strong Python; familiarity with AWS/GCP, containers, CI/CD; experience building infrastructure or data systems; authorized to intern in the U.S.
Senior Data & ML Infrastructure Engineer (Xora Portfolio Company)
Singapore or United States
HybridFull Time
Xora Innovation: Singapore-based private venture capital platform of Temasek investing in and building early-stage AI and deep-tech startups.
6+ YOEBachelor's or Master's in CS or related,6+ years building production software, strong Python, experience with large-scale data systems, ML infra, containers and orchestration, workflow orchestrators, and telemetry/monitoring.
Python, Airflow, Dagster, Flyte, Temporal, Docker, Kubernetes, Prometheus, Grafana, OpenTelemetry, Great Expectations, Evidently, MLflow, Weights & Biases, Ray Serve, KServe, Kubeflow, Atompack, ASE
Physical Intelligence: AI robotics developing foundation models and learning algorithms for robots and physically actuated devices.
Strong software engineering fundamentals; ML training infrastructure experience; large-scale JAX or PyTorch training; distributed systems, cloud workloads, performance optimization, and cross-functional communication experience.
Docker: Privately held container application platform helping developers build, share, and run applications.
5+ YOE5+ years applied ML/AI experience, 4+ years software engineering, experience with LLM-based systems, model lifecycle and ML infrastructure, bachelor's in CS/Engineering or equivalent, strong communication and mentoring skills.
Senior / Principal Infrastructure Engineer - ML Platform
San Mateo, California, United States
$279k-$345k/yrHybridFull Time
RobloxNYSE: RBLX: Global platform for user-created immersive digital experiences.
6+ YOE6+ years experience building scalable infrastructure; deep Kubernetes and Terraform experience; familiarity with AWS/GCP, Docker, CI/CD; bachelor's degree or equivalent practical experience.
SumerSports: AI-powered football intelligence serving professional and collegiate teams with scouting, roster, and performance analytics.
4+ YOERequires 4+ years in ML platform, DevOps, or infrastructure engineering; Kubernetes, CI/CD, containers, cloud, GPU clusters, Python, infrastructure as code, automation, observability, and production ML systems experience.
PhysicsX: Private physics-AI software helping industrial engineering and manufacturing teams design and optimize hardware.
5+ YOERequires 5+ years building and operating ML infrastructure at scale, distributed training expertise, Linux and networking fundamentals, Kubernetes and SLURM, Python, ML frameworks, and cloud GPU infrastructure experience.