496 ml infrastructure engineer jobs at 200 companies in San Bruno, CA
1mo
Save
Mark Applied
Hide
1mo
ML Infrastructure Engineer
Palo Alto, California, United States
$180k-$440k/yrOnsiteFull Time
xAI: Artificial intelligence research and development.
2+ YOE2+ years building large-scale production systems or ML infrastructure; degree in CS or related field or equivalent experience; strong Python and compiled-language skills; experience with GPU and distributed systems.
General MotorsNYSE: GM: Designing, building, and selling the world's best vehicles.
5+ YOE5+ years building large-scale distributed or ML systems; strong APIs and cloud infrastructure experience; expertise in ML lifecycle and MLOps; coding in Python or C++; BS/MS/PhD in CS/Math or equivalent experience.
Clera: AI-powered talent agent matching candidates to startup roles.
5+ YOE5+ years building production ML inference or model-serving systems. Requires scalable distributed systems, Docker, Kubernetes, cloud experience, observability tooling, and proficiency in Python, Go, Rust, C++, or Java.
Spellbrush: Generative AI and anime game studio making anime illustrations and video games for artists and players.
Experienced HPC/ML infrastructure engineer with Linux sysadmin skills, cluster bring-up and operations experience, familiarity with SLURM and parallel filesystems, networking and datacenter hardware handling.
General Legal: AI-native law firm serving growing companies with flat-fee commercial, corporate, and employment legal services.
5+ YOERequires 5+ years in software, machine learning, or infrastructure engineering; Python proficiency; production ML infrastructure, cloud, distributed systems, containers, independent system design, and a Computer Science degree or equivalent.
Nuro: Private U.S. autonomous-driving technology developing AI systems for automakers and mobility providers.
3+ YOE3+ years in ML infrastructure/backend platform or distributed systems. Experience with Terraform/Pulumi/Crossplane, Kubernetes/Ray/Slurm/Volcano schedulers, Apache Spark/Beam, feature stores (Feast/Hopsworks/Redis), and systems design for HPC.
Gridmatic: AI-powered energy supplying and optimizing clean electricity for commercial and industrial customers.
Significant experience building and operating production cloud infrastructure (GCP/AWS/Azure), Kubernetes (GKE), Terraform, workflow orchestration, Python and systems language experience; strong distributed systems and cloud networking skills.
Sr./Staff ML Infrastructure Engineer, Compute (TPU Scheduling) - Foundation Model
Cupertino, California, United States
OnsiteFull Time
AppleNASDAQ: AAPL: Designing and manufacturing consumer electronics, software, and digital services.
Experience building schedulers, resource managers, or orchestration systems for distributed workloads; experience with TPU/GPU accelerator infrastructure, distributed ML training/inference, and frameworks such as JAX, PyTorch, TensorFlow, Ray, Pathways; MS/PhD preferred.
Echo Neurotechnologies: Private San Francis neurotechnology startup developing brain-computer interfaces that restore communication for people with severe disabilities.
5+ YOEBachelor's in CS/EE or related,5+ years software or systems ML experience,proficient Python and PyTorch,distributed-training and large-scale data pipeline experience,excellent communication.
Gritt Robotics: Private AI-powered construction robotics automating labor-intensive infrastructure work for construction crews.
4+ YOE4+ years deploying high-performance ML pipelines; proficient in Python; comfortable with C++/Go; PyTorch; cloud platforms (AWS, GCP, Azure); Docker/Kubernetes/Airflow; Parquet/HDF5/TFRecord; US work authorization.
Mountain View or San Francisco or Kirkland or New York City
$175k-$215k/yrHybridFull Time
Waymo: Autonomous driving technology and robotaxi service provider.
Master's degree or equivalent practical experience; Python and C++ proficiency; modern deep learning framework familiarity; experience with large-scale data pipelines or ML infrastructure.
Docker: Privately held container application platform helping developers build, share, and run applications.
5+ YOE5+ years applied ML/AI experience, 4+ years software engineering, experience with LLM-based systems, model lifecycle and ML infrastructure, bachelor's in CS/Engineering or equivalent, strong communication and mentoring skills.
Senior / Principal Infrastructure Engineer - ML Platform
San Mateo, California, United States
$279k-$345k/yrHybridFull Time
RobloxNYSE: RBLX: Global platform for user-created immersive digital experiences.
6+ YOE6+ years experience building scalable infrastructure; deep Kubernetes and Terraform experience; familiarity with AWS/GCP, Docker, CI/CD; bachelor's degree or equivalent practical experience.
Skydio: American autonomous-drone manufacturer serving public safety, government, utility, and enterprise customers.
Experience with data engineering, cloud ML platforms, ML Ops, large-scale data pipelines, model training/deployment, and security/compliance in ML infrastructure; strong collaboration.
Cloud platforms, Containerization, ML Ops, Databases, Data pipelines
MakerMaker: Private U.S. AI research building autonomous research agents for recursive self-improvement.
6+ YOESenior ML engineer with 6+ years building production-grade ML systems; strong Python; distributed systems experience; familiar with Ray, Kubernetes, and experimentation infrastructure.
Sciforium: AI infrastructure building multimodal models and high-efficiency serving software for developers and teams.
5+ YOE5+ years in systems or infrastructure engineering with GPU, HPC, or ML infrastructure experience; technical bachelor's or master's degree; Linux, Kubernetes, schedulers, configuration management, Python, Bash, containers, GPUs, and RDMA expertise.
Cohesity: AI-powered data security and management solutions provider.
10+ YOERequires 10+ years software engineering, 8+ years distributed backend systems, 3+ years LLM/RAG or ML infrastructure, production Kubernetes experience, technical leadership, and a bachelor's degree in a technical field.
San Francisco or Los Angeles or Denver or Austin or Chicago or New York City or Canada or Seattle or Santa Barbara or San Diego or Toronto
$152k-$228k/yrRemoteFull Time
Invoca: Privately held SaaS platform helping marketing, commerce, and contact-center teams convert customer conversations into revenue.
5+ YOE5+ years of ML engineering experience; advanced Python, PyTorch, and deep learning; production NLP model deployment; SLM/LLM fine-tuning; inference infrastructure; production APIs; MLOps and model monitoring.