302 ml infrastructure engineer jobs at 100 companies in Aptos, CA
2w
Save
Mark Applied
Hide
2w
ML Infrastructure Engineer
Palo Alto, California, United States
$180k-$440k/yrOnsiteFull Time
xAI: Develops advanced artificial intelligence systems to understand the universe.
2+ YOE2+ years building large-scale production systems or ML infrastructure; degree in CS or related field or equivalent experience; strong Python and compiled-language skills; experience with GPU and distributed systems.
General MotorsNYSE: GM: Manufactures and sells automobiles and automotive parts globally.
5+ YOE5+ years building large-scale distributed or ML systems; strong APIs and cloud infrastructure experience; expertise in ML lifecycle and MLOps; coding in Python or C++; BS/MS/PhD in CS/Math or equivalent experience.
Nuro: Builds autonomous driving software and electric delivery robots.
3+ YOE3+ years in ML infrastructure/backend platform or distributed systems. Experience with Terraform/Pulumi/Crossplane, Kubernetes/Ray/Slurm/Volcano schedulers, Apache Spark/Beam, feature stores (Feast/Hopsworks/Redis), and systems design for HPC.
Gridmatic: AI-powered platform for optimizing energy trading and battery storage.
Significant experience building and operating production cloud infrastructure (GCP/AWS/Azure), Kubernetes (GKE), Terraform, workflow orchestration, Python and systems language experience; strong distributed systems and cloud networking skills.
Sr./Staff ML Infrastructure Engineer, Compute (TPU Scheduling) - Foundation Model
Cupertino, California, United States
OnsiteFull Time
AppleNASDAQ: AAPL: Designs and sells consumer electronics, software, and online services.
Experience building schedulers, resource managers, or orchestration systems for distributed workloads; experience with TPU/GPU accelerator infrastructure, distributed ML training/inference, and frameworks such as JAX, PyTorch, TensorFlow, Ray, Pathways; MS/PhD preferred.
Gritt Robotics: Developing autonomous robots to automate outdoor construction tasks.
4+ YOE4+ years deploying high-performance ML pipelines; proficient in Python; comfortable with C++/Go; PyTorch; cloud platforms (AWS, GCP, Azure); Docker/Kubernetes/Airflow; Parquet/HDF5/TFRecord; US work authorization.
MaxInsights: Provides robot data collection for physical AI development.
Experience building production ML training and deployment systems, strong software engineering and infra fundamentals, PyTorch experience, HPC/GPU knowledge, and strong communication and product sense.
Mountain View or San Francisco or Kirkland or New York City
$175k-$215k/yrHybridFull Time
Waymo: Autonomous driving technology for ride-hailing and logistics.
Master's degree or equivalent practical experience; Python and C++ proficiency; modern deep learning framework familiarity; experience with large-scale data pipelines or ML infrastructure.
Docker: Provides a platform for building, sharing, and running containerized applications.
5+ YOE5+ years applied ML/AI experience, 4+ years software engineering, experience with LLM-based systems, model lifecycle and ML infrastructure, bachelor's in CS/Engineering or equivalent, strong communication and mentoring skills.
Senior / Principal Infrastructure Engineer - ML Platform
San Mateo, California, United States
$279k-$345k/yrHybridFull Time
RobloxNYSE: RBLX: Platform for creating and playing user-generated 3D digital experiences.
6+ YOE6+ years experience building scalable infrastructure; deep Kubernetes and Terraform experience; familiarity with AWS/GCP, Docker, CI/CD; bachelor's degree or equivalent practical experience.
Skydio: Develops autonomous AI drones for defense and industrial inspection.
Experience with data engineering, cloud ML platforms, ML Ops, large-scale data pipelines, model training/deployment, and security/compliance in ML infrastructure; strong collaboration.
Cloud platforms, Containerization, ML Ops, Databases, Data pipelines
ML Systems Research Engineer, RL / Inference / Agent Systems
Santa Clara, California, United States
HybridFull Time
AMDNASDAQ: AMD: Designs and manufactures computer processors and graphics technology.
Experienced ML systems engineer with strong Python and ML framework skills, experience in RL/inference systems, distributed experimentation, and GPU/infrastructure workflows; advanced degree preferred.
Python, PyTorch, JAX, TensorFlow, Kubernetes, Ray, Slurm, ROCm, HIP, CUDA
Quince: Sells high-quality apparel and home goods at accessible prices.
8+ YOE8+ years industry experience with 4+ years in ML infrastructure/MLOps. Experience designing production ML platforms, cloud-native infra (AWS), Kubernetes, IaC, distributed training, feature stores, and cost/compute optimization.
NVIDIANASDAQ: NVDA: Designs graphics processing units and artificial intelligence hardware.
8+ YOE8+ years building scalable cloud infrastructure from concept to production; background in AI/ML or data analytics; strong OOP skills in Java or Go/Golang; experience with Kubernetes and distributed systems; Bachelor's or equivalent.
Rhoda AI: Developing generalist robotic intelligence for real-world industrial automation.
3+ YOE3+ years in ML infrastructure, MLOps, or distributed systems; strong Kubernetes; GPU orchestration; cloud/hybrid infra; ML frameworks; debugging ownership.
2+ YOEBachelor's degree or equivalent practical experience, 2+ years programming in Python or C++, and experience with ML infrastructure and a specialized ML area.
PayPalNASDAQ: PYPL: Digital platform for sending money and processing online payments.
10+ YOE10+ years relevant experience and a Bachelor’s degree (or equivalent); deep expertise in databases, data pipelines, messaging, caching, performance and resilience engineering; experience with real-time analytics and AI/ML infrastructure; strong technical leadership and mentoring.
CoupangNYSE: CPNG: Provides an end-to-end e-commerce and logistics network.
5+ YOEBachelor's in CS/EE/math,5+ years applied ML experience,proficiency in Python/Java,experience with big data pipelines,ML frameworks,cloud platforms,and building scalable low-latency services.
GuidewireNYSE: GWRE: Provides a software platform for property and casualty insurers.
10+ YOE10+ years software engineering; 5+ years ML platforms/infrastructure; distributed systems; Python/Go/Java; Docker/Kubernetes; MLOps tools; cloud experience; knowledge of ML models.
ML Systems Engineer, Large-Scale Model Training & RL Infrastructure
Palo Alto, California, United States
$195k-$262k/yrOnsiteFull Time
NebiusNasdaq: NBIS: Builds cloud infrastructure and software for artificial intelligence development.
Strong Python and PyTorch skills, hands-on distributed model training and GPU cluster experience, debugging across NCCL/CUDA/PyTorch/Ray, and quantitative reasoning about throughput, utilization, memory, and cost.