272 ml infrastructure engineer jobs at 80 companies in Soquel, CA

1w
Save
Mark Applied
Hide
ML Infrastructure Engineer
Palo Alto, California, United States
$180k-$440k/yr OnsiteFull Time
xAI
xAI: Develops advanced artificial intelligence systems to understand the universe.
2+ YOE2+ years building large-scale production systems or ML infrastructure; degree in CS or related field or equivalent experience; strong Python and compiled-language skills; experience with GPU and distributed systems.
Python, C++, Rust, JAX, PyTorch, NVIDIA drivers, CUDA, Linux, Slurm, Puppet, Ansible
1mo
Save
Mark Applied
Hide
Staff ML Infrastructure Engineer
Sunnyvale, California, United States
$189k-$291k/yr HybridFull Time
General Motors
General MotorsNYSE: GM: Manufactures and sells automobiles and automotive parts globally.
5+ YOE5+ years building large-scale distributed or ML systems; strong APIs and cloud infrastructure experience; expertise in ML lifecycle and MLOps; coding in Python or C++; BS/MS/PhD in CS/Math or equivalent experience.
Python, C++, PyTorch, TensorFlow, Bazel, Buck, Blaze, CMake, Docker, Kubernetes
1mo
Save
Mark Applied
Hide
Software Engineer, ML Infrastructure
Mountain View, California, United States
$160k-$241k/yr OnsiteFull Time
Nuro
Nuro: Builds autonomous driving software and electric delivery robots.
3+ YOE3+ years in ML infrastructure/backend platform or distributed systems. Experience with Terraform/Pulumi/Crossplane, Kubernetes/Ray/Slurm/Volcano schedulers, Apache Spark/Beam, feature stores (Feast/Hopsworks/Redis), and systems design for HPC.
Terraform, Pulumi, Crossplane, Kubernetes, KubeRay, Ray, Slurm, Volcano, Apache Spark, Apache Beam, Feast, Hopsworks, Redis, Lustre, Ceph, NVMe, AWS, GCP, Azure, Kubeflow, CNCF
1w
Save
Mark Applied
Hide
Sr./Staff ML Infrastructure Engineer, Compute (TPU Scheduling) - Foundation Model
Cupertino, California, United States
OnsiteFull Time
Apple
AppleNASDAQ: AAPL: Designs and sells consumer electronics, software, and online services.
Experience building schedulers, resource managers, or orchestration systems for distributed workloads; experience with TPU/GPU accelerator infrastructure, distributed ML training/inference, and frameworks such as JAX, PyTorch, TensorFlow, Ray, Pathways; MS/PhD preferred.
TPU, GPU, JAX, PyTorch, TensorFlow, Ray, Pathways
2mo
Save
Mark Applied
Hide
ML & Cloud Infrastructure Engineer
Belmont, California, United States
OnsiteFull Time
Gritt Robotics
Gritt Robotics: Developing autonomous robots to automate outdoor construction tasks.
4+ YOE4+ years deploying high-performance ML pipelines; proficient in Python; comfortable with C++/Go; PyTorch; cloud platforms (AWS, GCP, Azure); Docker/Kubernetes/Airflow; Parquet/HDF5/TFRecord; US work authorization.
Python, C++, Go, PyTorch, Parquet, HDF5, TFRecord, AWS, GCP, Azure, Docker, Kubernetes, Airflow
2w
Save
Mark Applied
Hide
Senior ML Infra Engineer
Santa Clara, California, United States
OnsiteFull Time
MaxInsights
MaxInsights: Provides robot data collection for physical AI development.
Experience building production ML training and deployment systems, strong software engineering and infra fundamentals, PyTorch experience, HPC/GPU knowledge, and strong communication and product sense.
Python, PyTorch, Docker, CI/CD, GPUs, CUDA, vLLM, TensorRT, Triton
1mo
Save
Mark Applied
Hide
ML Engineer
Palo Alto or Seattle or Paris
$139k-$226k/yr RemoteFull Time
Docker
Docker: Provides a platform for building, sharing, and running containerized applications.
5+ YOE5+ years applied ML/AI experience, 4+ years software engineering, experience with LLM-based systems, model lifecycle and ML infrastructure, bachelor's in CS/Engineering or equivalent, strong communication and mentoring skills.
Docker Desktop, Docker Hub, Docker Scout, LLM, MCP, Agentic Platform
3w
Save
Mark Applied
Hide
Senior / Principal Infrastructure Engineer - ML Platform
San Mateo, California, United States
$279k-$345k/yr HybridFull Time
Roblox
RobloxNYSE: RBLX: Platform for creating and playing user-generated 3D digital experiences.
6+ YOE6+ years experience building scalable infrastructure; deep Kubernetes and Terraform experience; familiarity with AWS/GCP, Docker, CI/CD; bachelor's degree or equivalent practical experience.
Kubernetes, Terraform, AWS, GCP, Docker, CI/CD
2mo
Save
Mark Applied
Hide
Autonomy Engineer - ML & DL Infrastructure
San Mateo, California, United States
$170k-$278k/yr HybridFull Time
Skydio
Skydio: Develops autonomous AI drones for defense and industrial inspection.
Experience with data engineering, cloud ML platforms, ML Ops, large-scale data pipelines, model training/deployment, and security/compliance in ML infrastructure; strong collaboration.
Cloud platforms, Containerization, ML Ops, Databases, Data pipelines
5d
Save
Mark Applied
Hide
ML Systems Research Engineer, RL / Inference / Agent Systems
Santa Clara, California, United States
HybridFull Time
AMD
AMDNASDAQ: AMD: Designs and manufactures computer processors and graphics technology.
Experienced ML systems engineer with strong Python and ML framework skills, experience in RL/inference systems, distributed experimentation, and GPU/infrastructure workflows; advanced degree preferred.
Python, PyTorch, JAX, TensorFlow, Kubernetes, Ray, Slurm, ROCm, HIP, CUDA
1w
Save
Mark Applied
Hide
Staff Engineer - ML Infra / MLOps
Palo Alto, California, United States
$218k-$285k/yr OnsiteFull Time
Quince
Quince: Sells high-quality apparel and home goods at accessible prices.
8+ YOE8+ years industry experience with 4+ years in ML infrastructure/MLOps. Experience designing production ML platforms, cloud-native infra (AWS), Kubernetes, IaC, distributed training, feature stores, and cost/compute optimization.
AWS, Kubernetes (EKS), Docker, Terraform, Pulumi, PyTorch, TensorFlow, Kubeflow, SageMaker, Spark, Flink, Kafka
2mo
Save
Mark Applied
Hide
Senior Machine Learning Infrastructure Engineer, Simulation
Mountain View or San Francisco
$213k-$263k/yr OnsiteFull Time
Waymo
Waymo: Autonomous driving technology for ride-hailing and logistics.
5+ YOE5+ years of professional software engineering, with 3+ years in ML infrastructure; BS in CS/Robotics or related field.
DeepSpeed, PyTorch, TensorFlow, ML infrastructure
2mo
Save
Mark Applied
Hide
Inference Infrastructure Engineer
Palo Alto, California, United States
OnsiteFull Time
Rhoda AI
Rhoda AI: Developing generalist robotic intelligence for real-world industrial automation.
3+ YOE3+ years in ML infrastructure, MLOps, or distributed systems; strong Kubernetes; GPU orchestration; cloud/hybrid infra; ML frameworks; debugging ownership.
Kubernetes, Triton, Ray Serve, TorchServe, PyTorch, JAX, AWS, GCP, Kafka, gRPC, NATS, SLURM, Ray, Volcano
1mo
Save
Mark Applied
Hide
Principal Engineer, Cloud Data Infrastructure
San Jose or Austin
$218k-$324k/yr HybridFull Time
PayPal
PayPalNASDAQ: PYPL: Digital platform for sending money and processing online payments.
10+ YOE10+ years relevant experience and a Bachelor’s degree (or equivalent); deep expertise in databases, data pipelines, messaging, caching, performance and resilience engineering; experience with real-time analytics and AI/ML infrastructure; strong technical leadership and mentoring.
2mo
Save
Mark Applied
Hide
Senior AI/ML Platform Engineer
San Mateo, California, United States
$148k-$247k/yr HybridFull Time
Guidewire
GuidewireNYSE: GWRE: Provides a software platform for property and casualty insurers.
10+ YOE10+ years software engineering; 5+ years ML platforms/infrastructure; distributed systems; Python/Go/Java; Docker/Kubernetes; MLOps tools; cloud experience; knowledge of ML models.
Python, Go, Java, Docker, Kubernetes, MLflow, Kubeflow, SageMaker, Vertex AI, Databricks, AWS, GCP, Azure
1w
Save
Mark Applied
Hide
ML Systems Engineer, Large-Scale Model Training & RL Infrastructure
Palo Alto, California, United States
$195k-$262k/yr OnsiteFull Time
Nebius
NebiusNasdaq: NBIS: Builds cloud infrastructure and software for artificial intelligence development.
Strong Python and PyTorch skills, hands-on distributed model training and GPU cluster experience, debugging across NCCL/CUDA/PyTorch/Ray, and quantitative reasoning about throughput, utilization, memory, and cost.
Python, PyTorch, Megatron-LM, DeepSpeed, PyTorch FSDP/DTensor, Ray, verl, slime, AReaL, OpenRLHF, NCCL, CUDA, Triton, Nsight, InfiniBand, RDMA, RoCE, Slurm, Kubernetes
2mo
Save
Mark Applied
Hide
Sr. ML Ops Engineer
Mountain View or United States
HybridFull Time
Corvus Robotics
Corvus Robotics: Fully autonomous drones for automated warehouse inventory management.
2+ YOE2-3 years building production ML infrastructure; experience with distributed data pipelines; understanding data flow from raw to trained models; ability to build from scratch or contribute to infra; thrive in ambiguous startup environment.
Kubeflow, SLURM, S3, data pipelines, distributed training
6d
Save
Mark Applied
Hide
Staff AI/ML Software Engineer, YouTube Ads Creative Foundational Infrastructure
Mountain View, California, United States
$207k-$300k/yr OnsiteFull Time
Google
GoogleNASDAQ: GOOGL: Provides online search, advertising, cloud computing, and consumer electronics.
8+ YOEBachelor's degree or equivalent,8+ years software development,5+ years product launches,experience with large-scale distributed systems and ML infrastructure,EMR not mentioned,TensorFlow/PyTorch proficiency preferred.
TensorFlow, PyTorch
1mo
Save
Mark Applied
Hide
Distinguished Engineer - Wireless Infrastructure
Santa Clara or New Hampshire or Massachusetts
$320k-$489k/yr HybridFull Time
NVIDIA
NVIDIANASDAQ: NVDA: Designs GPU-accelerated computing and artificial intelligence hardware.
20+ YOEMS/PhD in CS/EE or equivalent experience, 20+ years industry experience, deep expertise in communication systems, RAN algorithms, AI/ML for RAN, and programming in Python, C/C++, Matlab.
Python, C/C++, Matlab
1mo
Save
Mark Applied
Hide
Staff ML Infra Engineer, Search & Discovery
Mountain View, California, United States
$174k-$299k/yr OnsiteFull Time
Coupang
CoupangNYSE: CPNG: Provides online retail, grocery delivery, and video streaming services.
5+ YOEBachelor's in CS/EE/math/stats, 5+ years applied ML experience, production ML systems experience, proficiency in Python/Java, experience with big data, ML frameworks, and cloud platforms.
Python, Java, Hadoop, Hive, Presto, Spark, Scala, Apache Airflow, TensorFlow, PyTorch, AWS, GCP