254 ml infrastructure engineer jobs at 70 companies in Gilroy, CA
1w
Save
Mark Applied
Hide
1w
ML Infrastructure Engineer
Palo Alto, California, United States
$180k-$440k/yrOnsiteFull Time
xAI: Develops advanced artificial intelligence systems to understand the universe.
2+ YOE2+ years building large-scale production systems or ML infrastructure; degree in CS or related field or equivalent experience; strong Python and compiled-language skills; experience with GPU and distributed systems.
General MotorsNYSE: GM: Manufactures and sells automobiles and automotive parts globally.
5+ YOE5+ years building large-scale distributed or ML systems; strong APIs and cloud infrastructure experience; expertise in ML lifecycle and MLOps; coding in Python or C++; BS/MS/PhD in CS/Math or equivalent experience.
Nuro: Builds autonomous driving software and electric delivery robots.
3+ YOE3+ years in ML infrastructure/backend platform or distributed systems. Experience with Terraform/Pulumi/Crossplane, Kubernetes/Ray/Slurm/Volcano schedulers, Apache Spark/Beam, feature stores (Feast/Hopsworks/Redis), and systems design for HPC.
Sr./Staff ML Infrastructure Engineer, Compute (TPU Scheduling) - Foundation Model
Cupertino, California, United States
OnsiteFull Time
AppleNASDAQ: AAPL: Designs and sells consumer electronics, software, and online services.
Experience building schedulers, resource managers, or orchestration systems for distributed workloads; experience with TPU/GPU accelerator infrastructure, distributed ML training/inference, and frameworks such as JAX, PyTorch, TensorFlow, Ray, Pathways; MS/PhD preferred.
MaxInsights: Provides robot data collection for physical AI development.
Experience building production ML training and deployment systems, strong software engineering and infra fundamentals, PyTorch experience, HPC/GPU knowledge, and strong communication and product sense.
Docker: Provides a platform for building, sharing, and running containerized applications.
5+ YOE5+ years applied ML/AI experience, 4+ years software engineering, experience with LLM-based systems, model lifecycle and ML infrastructure, bachelor's in CS/Engineering or equivalent, strong communication and mentoring skills.
ML Systems Research Engineer, RL / Inference / Agent Systems
Santa Clara, California, United States
HybridFull Time
AMDNASDAQ: AMD: Designs and manufactures computer processors and graphics technology.
Experienced ML systems engineer with strong Python and ML framework skills, experience in RL/inference systems, distributed experimentation, and GPU/infrastructure workflows; advanced degree preferred.
Python, PyTorch, JAX, TensorFlow, Kubernetes, Ray, Slurm, ROCm, HIP, CUDA
Quince: Sells high-quality apparel and home goods at accessible prices.
8+ YOE8+ years industry experience with 4+ years in ML infrastructure/MLOps. Experience designing production ML platforms, cloud-native infra (AWS), Kubernetes, IaC, distributed training, feature stores, and cost/compute optimization.
Rhoda AI: Developing generalist robotic intelligence for real-world industrial automation.
3+ YOE3+ years in ML infrastructure, MLOps, or distributed systems; strong Kubernetes; GPU orchestration; cloud/hybrid infra; ML frameworks; debugging ownership.
PayPalNASDAQ: PYPL: Digital platform for sending money and processing online payments.
10+ YOE10+ years relevant experience and a Bachelor’s degree (or equivalent); deep expertise in databases, data pipelines, messaging, caching, performance and resilience engineering; experience with real-time analytics and AI/ML infrastructure; strong technical leadership and mentoring.
ML Systems Engineer, Large-Scale Model Training & RL Infrastructure
Palo Alto, California, United States
$195k-$262k/yrOnsiteFull Time
NebiusNasdaq: NBIS: Builds cloud infrastructure and software for artificial intelligence development.
Strong Python and PyTorch skills, hands-on distributed model training and GPU cluster experience, debugging across NCCL/CUDA/PyTorch/Ray, and quantitative reasoning about throughput, utilization, memory, and cost.
Corvus Robotics: Fully autonomous drones for automated warehouse inventory management.
2+ YOE2-3 years building production ML infrastructure; experience with distributed data pipelines; understanding data flow from raw to trained models; ability to build from scratch or contribute to infra; thrive in ambiguous startup environment.
Kubeflow, SLURM, S3, data pipelines, distributed training
8+ YOEBachelor's degree or equivalent,8+ years software development,5+ years product launches,experience with large-scale distributed systems and ML infrastructure,EMR not mentioned,TensorFlow/PyTorch proficiency preferred.
NVIDIANASDAQ: NVDA: Designs GPU-accelerated computing and artificial intelligence hardware.
20+ YOEMS/PhD in CS/EE or equivalent experience, 20+ years industry experience, deep expertise in communication systems, RAN algorithms, AI/ML for RAN, and programming in Python, C/C++, Matlab.
CoupangNYSE: CPNG: Provides online retail, grocery delivery, and video streaming services.
5+ YOEBachelor's in CS/EE/math/stats, 5+ years applied ML experience, production ML systems experience, proficiency in Python/Java, experience with big data, ML frameworks, and cloud platforms.
TikTok: Global short-form video hosting and social media platform.
3+ YOE3+ years building scalable ML systems, strong CS fundamentals, coding skills, experience with causal inference/uplift/deep learning, project management and communication skills.
Senior AI Infrastructure Engineer - Model Training
Mountain View, California, United States
$190k-$260k/yrOnsiteFull Time
Kodiak RoboticsNASDAQ: KDK: Develops autonomous driving technology for commercial trucking and defense.
2+ YOEDegree in CS or related field,2+ years ML systems experience,expertise in distributed training,high-performance data pipelines,GPU performance and profiling,Python and PyTorch skills.
TargetNYSE: TGT: Operates a chain of general merchandise stores and supermarkets.
MS or equivalent preferred; extensive experience designing and operating large-scale cloud-native ML platforms, Kubernetes-based infrastructure, MLOps, model governance, observability, and platform automation.