254 ml infrastructure engineer jobs at 70 companies in Gilroy, CA

1w
Save
Mark Applied
Hide
ML Infrastructure Engineer
Palo Alto, California, United States
$180k-$440k/yr OnsiteFull Time
xAI
xAI: Develops advanced artificial intelligence systems to understand the universe.
2+ YOE2+ years building large-scale production systems or ML infrastructure; degree in CS or related field or equivalent experience; strong Python and compiled-language skills; experience with GPU and distributed systems.
Python, C++, Rust, JAX, PyTorch, NVIDIA drivers, CUDA, Linux, Slurm, Puppet, Ansible
1mo
Save
Mark Applied
Hide
Staff ML Infrastructure Engineer
Sunnyvale, California, United States
$189k-$291k/yr HybridFull Time
General Motors
General MotorsNYSE: GM: Manufactures and sells automobiles and automotive parts globally.
5+ YOE5+ years building large-scale distributed or ML systems; strong APIs and cloud infrastructure experience; expertise in ML lifecycle and MLOps; coding in Python or C++; BS/MS/PhD in CS/Math or equivalent experience.
Python, C++, PyTorch, TensorFlow, Bazel, Buck, Blaze, CMake, Docker, Kubernetes
1mo
Save
Mark Applied
Hide
Software Engineer, ML Infrastructure
Mountain View, California, United States
$160k-$241k/yr OnsiteFull Time
Nuro
Nuro: Builds autonomous driving software and electric delivery robots.
3+ YOE3+ years in ML infrastructure/backend platform or distributed systems. Experience with Terraform/Pulumi/Crossplane, Kubernetes/Ray/Slurm/Volcano schedulers, Apache Spark/Beam, feature stores (Feast/Hopsworks/Redis), and systems design for HPC.
Terraform, Pulumi, Crossplane, Kubernetes, KubeRay, Ray, Slurm, Volcano, Apache Spark, Apache Beam, Feast, Hopsworks, Redis, Lustre, Ceph, NVMe, AWS, GCP, Azure, Kubeflow, CNCF
1w
Save
Mark Applied
Hide
Sr./Staff ML Infrastructure Engineer, Compute (TPU Scheduling) - Foundation Model
Cupertino, California, United States
OnsiteFull Time
Apple
AppleNASDAQ: AAPL: Designs and sells consumer electronics, software, and online services.
Experience building schedulers, resource managers, or orchestration systems for distributed workloads; experience with TPU/GPU accelerator infrastructure, distributed ML training/inference, and frameworks such as JAX, PyTorch, TensorFlow, Ray, Pathways; MS/PhD preferred.
TPU, GPU, JAX, PyTorch, TensorFlow, Ray, Pathways
2w
Save
Mark Applied
Hide
Senior ML Infra Engineer
Santa Clara, California, United States
OnsiteFull Time
MaxInsights
MaxInsights: Provides robot data collection for physical AI development.
Experience building production ML training and deployment systems, strong software engineering and infra fundamentals, PyTorch experience, HPC/GPU knowledge, and strong communication and product sense.
Python, PyTorch, Docker, CI/CD, GPUs, CUDA, vLLM, TensorRT, Triton
1mo
Save
Mark Applied
Hide
ML Engineer
Palo Alto or Seattle or Paris
$139k-$226k/yr RemoteFull Time
Docker
Docker: Provides a platform for building, sharing, and running containerized applications.
5+ YOE5+ years applied ML/AI experience, 4+ years software engineering, experience with LLM-based systems, model lifecycle and ML infrastructure, bachelor's in CS/Engineering or equivalent, strong communication and mentoring skills.
Docker Desktop, Docker Hub, Docker Scout, LLM, MCP, Agentic Platform
3w
Save
Mark Applied
Hide
Sr. Mgr., ML Infrastructure, PV Personalization and Discovery
Sunnyvale or Seattle or New York
$242k-$328k/yr OnsiteFull Time
Amazon
AmazonNASDAQ: AMZN: Global online retail and cloud computing technology provider.
10+ YOE5+ Mgmt10+ years engineering experience, 5+ years managing engineering teams, expertise in ML infrastructure, retrieval/recommendation systems, LLMs/generative AI, partnering with applied scientists, experience delivering large-scale consumer software.
AWS
5d
Save
Mark Applied
Hide
ML Systems Research Engineer, RL / Inference / Agent Systems
Santa Clara, California, United States
HybridFull Time
AMD
AMDNASDAQ: AMD: Designs and manufactures computer processors and graphics technology.
Experienced ML systems engineer with strong Python and ML framework skills, experience in RL/inference systems, distributed experimentation, and GPU/infrastructure workflows; advanced degree preferred.
Python, PyTorch, JAX, TensorFlow, Kubernetes, Ray, Slurm, ROCm, HIP, CUDA
1w
Save
Mark Applied
Hide
Staff Engineer - ML Infra / MLOps
Palo Alto, California, United States
$218k-$285k/yr OnsiteFull Time
Quince
Quince: Sells high-quality apparel and home goods at accessible prices.
8+ YOE8+ years industry experience with 4+ years in ML infrastructure/MLOps. Experience designing production ML platforms, cloud-native infra (AWS), Kubernetes, IaC, distributed training, feature stores, and cost/compute optimization.
AWS, Kubernetes (EKS), Docker, Terraform, Pulumi, PyTorch, TensorFlow, Kubeflow, SageMaker, Spark, Flink, Kafka
2mo
Save
Mark Applied
Hide
Senior Machine Learning Infrastructure Engineer, Simulation
Mountain View or San Francisco
$213k-$263k/yr OnsiteFull Time
Waymo
Waymo: Autonomous driving technology for ride-hailing and logistics.
5+ YOE5+ years of professional software engineering, with 3+ years in ML infrastructure; BS in CS/Robotics or related field.
DeepSpeed, PyTorch, TensorFlow, ML infrastructure
2mo
Save
Mark Applied
Hide
Inference Infrastructure Engineer
Palo Alto, California, United States
OnsiteFull Time
Rhoda AI
Rhoda AI: Developing generalist robotic intelligence for real-world industrial automation.
3+ YOE3+ years in ML infrastructure, MLOps, or distributed systems; strong Kubernetes; GPU orchestration; cloud/hybrid infra; ML frameworks; debugging ownership.
Kubernetes, Triton, Ray Serve, TorchServe, PyTorch, JAX, AWS, GCP, Kafka, gRPC, NATS, SLURM, Ray, Volcano
1mo
Save
Mark Applied
Hide
Principal Engineer, Cloud Data Infrastructure
San Jose or Austin
$218k-$324k/yr HybridFull Time
PayPal
PayPalNASDAQ: PYPL: Digital platform for sending money and processing online payments.
10+ YOE10+ years relevant experience and a Bachelor’s degree (or equivalent); deep expertise in databases, data pipelines, messaging, caching, performance and resilience engineering; experience with real-time analytics and AI/ML infrastructure; strong technical leadership and mentoring.
1w
Save
Mark Applied
Hide
ML Systems Engineer, Large-Scale Model Training & RL Infrastructure
Palo Alto, California, United States
$195k-$262k/yr OnsiteFull Time
Nebius
NebiusNasdaq: NBIS: Builds cloud infrastructure and software for artificial intelligence development.
Strong Python and PyTorch skills, hands-on distributed model training and GPU cluster experience, debugging across NCCL/CUDA/PyTorch/Ray, and quantitative reasoning about throughput, utilization, memory, and cost.
Python, PyTorch, Megatron-LM, DeepSpeed, PyTorch FSDP/DTensor, Ray, verl, slime, AReaL, OpenRLHF, NCCL, CUDA, Triton, Nsight, InfiniBand, RDMA, RoCE, Slurm, Kubernetes
2mo
Save
Mark Applied
Hide
Sr. ML Ops Engineer
Mountain View or United States
HybridFull Time
Corvus Robotics
Corvus Robotics: Fully autonomous drones for automated warehouse inventory management.
2+ YOE2-3 years building production ML infrastructure; experience with distributed data pipelines; understanding data flow from raw to trained models; ability to build from scratch or contribute to infra; thrive in ambiguous startup environment.
Kubeflow, SLURM, S3, data pipelines, distributed training
6d
Save
Mark Applied
Hide
Staff AI/ML Software Engineer, YouTube Ads Creative Foundational Infrastructure
Mountain View, California, United States
$207k-$300k/yr OnsiteFull Time
Google
GoogleNASDAQ: GOOGL: Provides online search, advertising, cloud computing, and consumer electronics.
8+ YOEBachelor's degree or equivalent,8+ years software development,5+ years product launches,experience with large-scale distributed systems and ML infrastructure,EMR not mentioned,TensorFlow/PyTorch proficiency preferred.
TensorFlow, PyTorch
1mo
Save
Mark Applied
Hide
Distinguished Engineer - Wireless Infrastructure
Santa Clara or New Hampshire or Massachusetts
$320k-$489k/yr HybridFull Time
NVIDIA
NVIDIANASDAQ: NVDA: Designs GPU-accelerated computing and artificial intelligence hardware.
20+ YOEMS/PhD in CS/EE or equivalent experience, 20+ years industry experience, deep expertise in communication systems, RAN algorithms, AI/ML for RAN, and programming in Python, C/C++, Matlab.
Python, C/C++, Matlab
1mo
Save
Mark Applied
Hide
Staff ML Infra Engineer, Search & Discovery
Mountain View, California, United States
$174k-$299k/yr OnsiteFull Time
Coupang
CoupangNYSE: CPNG: Provides online retail, grocery delivery, and video streaming services.
5+ YOEBachelor's in CS/EE/math/stats, 5+ years applied ML experience, production ML systems experience, proficiency in Python/Java, experience with big data, ML frameworks, and cloud platforms.
Python, Java, Hadoop, Hive, Presto, Spark, Scala, Apache Airflow, TensorFlow, PyTorch, AWS, GCP
3w
Save
Mark Applied
Hide
Software Engineer, Ads ML Infrastructure
San Jose, California, United States
$156k-$317k/yr OnsiteFull Time
TikTok
TikTok: Global short-form video hosting and social media platform.
3+ YOE3+ years building scalable ML systems, strong CS fundamentals, coding skills, experience with causal inference/uplift/deep learning, project management and communication skills.
3w
Save
Mark Applied
Hide
Senior AI Infrastructure Engineer - Model Training
Mountain View, California, United States
$190k-$260k/yr OnsiteFull Time
Kodiak Robotics
Kodiak RoboticsNASDAQ: KDK: Develops autonomous driving technology for commercial trucking and defense.
2+ YOEDegree in CS or related field,2+ years ML systems experience,expertise in distributed training,high-performance data pipelines,GPU performance and profiling,Python and PyTorch skills.
PyTorch, PyTorch DDP/FSDP, DeepSpeed, Megatron, NCCL, WebDataset, MosaicML Streaming, MDS, Nsight, PyTorch Profiler, Python, C++, CUDA, Triton, NVLink, InfiniBand
2w
Save
Mark Applied
Hide
Principal Engineer - AI /ML Platform
Brooklyn Park or Sunnyvale
$168k-$356k/yr HybridFull Time
Target
TargetNYSE: TGT: Operates a chain of general merchandise stores and supermarkets.
MS or equivalent preferred; extensive experience designing and operating large-scale cloud-native ML platforms, Kubernetes-based infrastructure, MLOps, model governance, observability, and platform automation.
Vertex AI, Kubeflow, MLflow, Terraform, GitOps, Kubernetes, CI/CD