302 ml infrastructure engineer jobs at 100 companies in Aptos, CA

2w
Save
Mark Applied
Hide
ML Infrastructure Engineer
Palo Alto, California, United States
$180k-$440k/yr OnsiteFull Time
xAI
xAI: Develops advanced artificial intelligence systems to understand the universe.
2+ YOE2+ years building large-scale production systems or ML infrastructure; degree in CS or related field or equivalent experience; strong Python and compiled-language skills; experience with GPU and distributed systems.
Python, C++, Rust, JAX, PyTorch, NVIDIA drivers, CUDA, Linux, Slurm, Puppet, Ansible
1mo
Save
Mark Applied
Hide
Staff ML Infrastructure Engineer
Sunnyvale, California, United States
$189k-$291k/yr HybridFull Time
General Motors
General MotorsNYSE: GM: Manufactures and sells automobiles and automotive parts globally.
5+ YOE5+ years building large-scale distributed or ML systems; strong APIs and cloud infrastructure experience; expertise in ML lifecycle and MLOps; coding in Python or C++; BS/MS/PhD in CS/Math or equivalent experience.
Python, C++, PyTorch, TensorFlow, Bazel, Buck, Blaze, CMake, Docker, Kubernetes
1mo
Save
Mark Applied
Hide
Software Engineer, ML Infrastructure
Mountain View, California, United States
$160k-$241k/yr OnsiteFull Time
Nuro
Nuro: Builds autonomous driving software and electric delivery robots.
3+ YOE3+ years in ML infrastructure/backend platform or distributed systems. Experience with Terraform/Pulumi/Crossplane, Kubernetes/Ray/Slurm/Volcano schedulers, Apache Spark/Beam, feature stores (Feast/Hopsworks/Redis), and systems design for HPC.
Terraform, Pulumi, Crossplane, Kubernetes, KubeRay, Ray, Slurm, Volcano, Apache Spark, Apache Beam, Feast, Hopsworks, Redis, Lustre, Ceph, NVMe, AWS, GCP, Azure, Kubeflow, CNCF
1w
Save
Mark Applied
Hide
Senior Software Engineer, ML Infrastructure
Cupertino or San Francisco or Houston
$209k-$235k/yr OnsiteFull Time
Gridmatic
Gridmatic: AI-powered platform for optimizing energy trading and battery storage.
Significant experience building and operating production cloud infrastructure (GCP/AWS/Azure), Kubernetes (GKE), Terraform, workflow orchestration, Python and systems language experience; strong distributed systems and cloud networking skills.
GCP, AWS, Azure, Kubernetes, GKE, Terraform, Flyte, Temporal, Airflow, Python, Go, C++, Java, Rust, Grafana, Google Cloud Monitoring
3w
Save
Mark Applied
Hide
Sr./Staff ML Infrastructure Engineer, Compute (TPU Scheduling) - Foundation Model
Cupertino, California, United States
OnsiteFull Time
Apple
AppleNASDAQ: AAPL: Designs and sells consumer electronics, software, and online services.
Experience building schedulers, resource managers, or orchestration systems for distributed workloads; experience with TPU/GPU accelerator infrastructure, distributed ML training/inference, and frameworks such as JAX, PyTorch, TensorFlow, Ray, Pathways; MS/PhD preferred.
TPU, GPU, JAX, PyTorch, TensorFlow, Ray, Pathways
2mo
Save
Mark Applied
Hide
ML & Cloud Infrastructure Engineer
Belmont, California, United States
OnsiteFull Time
Gritt Robotics
Gritt Robotics: Developing autonomous robots to automate outdoor construction tasks.
4+ YOE4+ years deploying high-performance ML pipelines; proficient in Python; comfortable with C++/Go; PyTorch; cloud platforms (AWS, GCP, Azure); Docker/Kubernetes/Airflow; Parquet/HDF5/TFRecord; US work authorization.
Python, C++, Go, PyTorch, Parquet, HDF5, TFRecord, AWS, GCP, Azure, Docker, Kubernetes, Airflow
3w
Save
Mark Applied
Hide
Senior ML Infra Engineer
Santa Clara, California, United States
OnsiteFull Time
MaxInsights
MaxInsights: Provides robot data collection for physical AI development.
Experience building production ML training and deployment systems, strong software engineering and infra fundamentals, PyTorch experience, HPC/GPU knowledge, and strong communication and product sense.
Python, PyTorch, Docker, CI/CD, GPUs, CUDA, vLLM, TensorRT, Triton
5d
Save
Mark Applied
Hide
ML Engineer, Foundation Model Infrastructure
Mountain View or San Francisco or Kirkland or New York City
$175k-$215k/yr HybridFull Time
Waymo
Waymo: Autonomous driving technology for ride-hailing and logistics.
Master's degree or equivalent practical experience; Python and C++ proficiency; modern deep learning framework familiarity; experience with large-scale data pipelines or ML infrastructure.
Python, C++, Flume, JAX, PyTorch, TensorFlow, Spark, Borg, Kubeflow
1mo
Save
Mark Applied
Hide
ML Engineer
Palo Alto or Seattle or Paris
$139k-$226k/yr RemoteFull Time
Docker
Docker: Provides a platform for building, sharing, and running containerized applications.
5+ YOE5+ years applied ML/AI experience, 4+ years software engineering, experience with LLM-based systems, model lifecycle and ML infrastructure, bachelor's in CS/Engineering or equivalent, strong communication and mentoring skills.
Docker Desktop, Docker Hub, Docker Scout, LLM, MCP, Agentic Platform
1mo
Save
Mark Applied
Hide
Senior / Principal Infrastructure Engineer - ML Platform
San Mateo, California, United States
$279k-$345k/yr HybridFull Time
Roblox
RobloxNYSE: RBLX: Platform for creating and playing user-generated 3D digital experiences.
6+ YOE6+ years experience building scalable infrastructure; deep Kubernetes and Terraform experience; familiarity with AWS/GCP, Docker, CI/CD; bachelor's degree or equivalent practical experience.
Kubernetes, Terraform, AWS, GCP, Docker, CI/CD
3mo
Save
Mark Applied
Hide
Autonomy Engineer - ML & DL Infrastructure
San Mateo, California, United States
$170k-$278k/yr HybridFull Time
Skydio
Skydio: Develops autonomous AI drones for defense and industrial inspection.
Experience with data engineering, cloud ML platforms, ML Ops, large-scale data pipelines, model training/deployment, and security/compliance in ML infrastructure; strong collaboration.
Cloud platforms, Containerization, ML Ops, Databases, Data pipelines
1w
Save
Mark Applied
Hide
ML Systems Research Engineer, RL / Inference / Agent Systems
Santa Clara, California, United States
HybridFull Time
AMD
AMDNASDAQ: AMD: Designs and manufactures computer processors and graphics technology.
Experienced ML systems engineer with strong Python and ML framework skills, experience in RL/inference systems, distributed experimentation, and GPU/infrastructure workflows; advanced degree preferred.
Python, PyTorch, JAX, TensorFlow, Kubernetes, Ray, Slurm, ROCm, HIP, CUDA
2w
Save
Mark Applied
Hide
Staff Engineer - ML Infra / MLOps
Palo Alto, California, United States
$218k-$285k/yr OnsiteFull Time
Quince
Quince: Sells high-quality apparel and home goods at accessible prices.
8+ YOE8+ years industry experience with 4+ years in ML infrastructure/MLOps. Experience designing production ML platforms, cloud-native infra (AWS), Kubernetes, IaC, distributed training, feature stores, and cost/compute optimization.
AWS, Kubernetes (EKS), Docker, Terraform, Pulumi, PyTorch, TensorFlow, Kubeflow, SageMaker, Spark, Flink, Kafka
1mo
Save
Mark Applied
Hide
Senior Infrastructure Engineer – Bazel Remote Execution
Santa Clara, California, United States
$184k-$357k/yr OnsiteFull Time
NVIDIA
NVIDIANASDAQ: NVDA: Designs graphics processing units and artificial intelligence hardware.
8+ YOE8+ years building scalable cloud infrastructure from concept to production; background in AI/ML or data analytics; strong OOP skills in Java or Go/Golang; experience with Kubernetes and distributed systems; Bachelor's or equivalent.
Bazel Remote Execution, Buildbarn, Buildfarm, Kubernetes, Java, Go, Golang
3mo
Save
Mark Applied
Hide
Inference Infrastructure Engineer
Palo Alto, California, United States
OnsiteFull Time
Rhoda AI
Rhoda AI: Developing generalist robotic intelligence for real-world industrial automation.
3+ YOE3+ years in ML infrastructure, MLOps, or distributed systems; strong Kubernetes; GPU orchestration; cloud/hybrid infra; ML frameworks; debugging ownership.
Kubernetes, Triton, Ray Serve, TorchServe, PyTorch, JAX, AWS, GCP, Kafka, gRPC, NATS, SLURM, Ray, Volcano
17h
Save
Mark Applied
Hide
Software Engineer III, AI/ML, AI and Infrastructure
Sunnyvale, California, United States
$147k-$210k/yr OnsiteFull Time
Google
GoogleNASDAQ: GOOGL: Provides online search, advertising, cloud computing, and consumer electronics.
2+ YOEBachelor's degree or equivalent practical experience, 2+ years programming in Python or C++, and experience with ML infrastructure and a specialized ML area.
Python, C++, Machine Learning (ML), Machine Learning Infrastructure, Vertex AI, TPUs
1mo
Save
Mark Applied
Hide
Principal Engineer, Cloud Data Infrastructure
San Jose or Austin
$218k-$324k/yr HybridFull Time
PayPal
PayPalNASDAQ: PYPL: Digital platform for sending money and processing online payments.
10+ YOE10+ years relevant experience and a Bachelor’s degree (or equivalent); deep expertise in databases, data pipelines, messaging, caching, performance and resilience engineering; experience with real-time analytics and AI/ML infrastructure; strong technical leadership and mentoring.
5d
Save
Mark Applied
Hide
Staff ML Infra Engineer
Mountain View, California, United States
$174k-$299k/yr OnsiteFull Time
Coupang
CoupangNYSE: CPNG: Provides an end-to-end e-commerce and logistics network.
5+ YOEBachelor's in CS/EE/math,5+ years applied ML experience,proficiency in Python/Java,experience with big data pipelines,ML frameworks,cloud platforms,and building scalable low-latency services.
Python, Java, Hadoop, Hive, Presto, Spark, Scala, Apache Airflow, TensorFlow, PyTorch, AWS, GCP
2mo
Save
Mark Applied
Hide
Senior AI/ML Platform Engineer
San Mateo, California, United States
$148k-$247k/yr HybridFull Time
Guidewire
GuidewireNYSE: GWRE: Provides a software platform for property and casualty insurers.
10+ YOE10+ years software engineering; 5+ years ML platforms/infrastructure; distributed systems; Python/Go/Java; Docker/Kubernetes; MLOps tools; cloud experience; knowledge of ML models.
Python, Go, Java, Docker, Kubernetes, MLflow, Kubeflow, SageMaker, Vertex AI, Databricks, AWS, GCP, Azure
2w
Save
Mark Applied
Hide
ML Systems Engineer, Large-Scale Model Training & RL Infrastructure
Palo Alto, California, United States
$195k-$262k/yr OnsiteFull Time
Nebius
NebiusNasdaq: NBIS: Builds cloud infrastructure and software for artificial intelligence development.
Strong Python and PyTorch skills, hands-on distributed model training and GPU cluster experience, debugging across NCCL/CUDA/PyTorch/Ray, and quantitative reasoning about throughput, utilization, memory, and cost.
Python, PyTorch, Megatron-LM, DeepSpeed, PyTorch FSDP/DTensor, Ray, verl, slime, AReaL, OpenRLHF, NCCL, CUDA, Triton, Nsight, InfiniBand, RDMA, RoCE, Slurm, Kubernetes