930 ml systems engineer jobs at 350 companies in Oakley, CA
2mo
Save
Mark Applied
Hide
2mo
ML Systems Engineer
San Jose or Santa Barbara
$150k-$350k/yrOnsiteFull Time
Alpha Design AI: AI-native EDA platform for semiconductor design and verification.
Experience with large-scale ML systems and GPU computing; strong Python and C++/CUDA skills; familiarity with vLLM, PyTorch, SGLang, Ray; experience deploying and optimizing LLMs, profiling and benchmarking inference.
Engram: Developing persistent memory layers for enterprise AI systems.
5+ YOE5+ years building training/inference systems; strong engineering skills; experience with ML frameworks, GPUs, distributed systems; bachelor's degree or equivalent experience.
ML Systems Research Engineer, RL / Inference / Agent Systems
Santa Clara, California, United States
HybridFull Time
AMDNASDAQ: AMD: Designs and manufactures computer processors and graphics technology.
Experienced ML systems engineer with strong Python and ML framework skills, experience in RL/inference systems, distributed experimentation, and GPU/infrastructure workflows; advanced degree preferred.
Python, PyTorch, JAX, TensorFlow, Kubernetes, Ray, Slurm, ROCm, HIP, CUDA
NVIDIANASDAQ: NVDA: Designs GPU-accelerated computing and artificial intelligence hardware.
12+ YOE12+ years building ML systems and software platforms; expert Python and PyTorch; experience with ML pipelines, evaluation, developer tooling, and LLM/agentic systems; BS/MS in CS, Engineering, or equivalent experience.
Docker: Provides a platform for building, sharing, and running containerized applications.
8+ YOE8+ years professional software engineering experience, 5+ years applied ML experience, bachelor's in CS/Engineering or equivalent, experience shipping ML systems, LLM/agent experience, on-call participation possible.
Physical Intelligence: Creating foundation models for general-purpose robot intelligence.
Strong software engineering fundamentals with experience building distributed systems, large-scale data pipelines, object storage and batch/streaming systems; ownership mindset and performance focus.
RivianNASDAQ: RIVN: Designs and manufactures electric vehicles and charging networks.
5+ YOE5+ years building and scaling ML solutions for auto-labeling and AV perception; strong Python, perception pipeline and system engineering experience; BS/MS/PhD in CS/Robotics/Electrical Engineering or related.
New York City or San Francisco or London or Sydney
$240k-$270k/yrOnsiteFull Time
Sigma Computing: Cloud-native analytics platform featuring a spreadsheet-style interface.
10+ YOEBachelor's in CS/Engineering/Mathematics, 10+ years building and deploying production AI/ML systems, expertise in ML/DL, full ML lifecycle experience, foundation-model adaptation experience.
MakerMaker: Small San Francis-based team building autonomous ML systems
6+ YOESenior ML engineer with 6+ years building production-grade ML systems; strong Python; distributed systems experience; familiar with Ray, Kubernetes, and experimentation infrastructure.
Sunnyvale or Austin or Detroit or Warren or Milford or Mountain View or United States
$296k-$424k/yrHybridFull Time
General MotorsNYSE: GM: Manufactures and sells automobiles and automotive parts globally.
MS or PhD in CS/Robotics/ML or related; experience leading technical teams delivering production ML systems; deep expertise in robotics planning/control, imitation or reinforcement learning, generative models, and large-scale ML; strong software engineering skills in Python and C++.
xAI: Develops advanced artificial intelligence systems to understand the universe.
2+ YOE2+ years building large-scale production systems or ML infrastructure; degree in CS or related field or equivalent experience; strong Python and compiled-language skills; experience with GPU and distributed systems.
Gap Inc.NYSE: GAP: Global specialty retailer of apparel and accessories.
10+ YOE10+ years building production ML systems; strong Python and software engineering; experience with ML frameworks, model serving, MLOps, cloud platforms, and distributed data processing.
MaxInsights: Provides robot data collection for physical AI development.
Experience building production ML training and deployment systems, strong software engineering and infra fundamentals, PyTorch experience, HPC/GPU knowledge, and strong communication and product sense.
ML Systems Engineer, Large-Scale Model Training & RL Infrastructure
Palo Alto, California, United States
$195k-$262k/yrOnsiteFull Time
NebiusNasdaq: NBIS: Builds cloud infrastructure and software for artificial intelligence development.
Strong Python and PyTorch skills, hands-on distributed model training and GPU cluster experience, debugging across NCCL/CUDA/PyTorch/Ray, and quantitative reasoning about throughput, utilization, memory, and cost.
Clera: AI talent agent matching professionals with high-growth startup roles
4+ YOE4+ years applied ML engineering in production, experience with LLMs/fine-tuning/RAG or large-scale recommender systems, strong Python and PyTorch or JAX skills, distributed training/GPU/inference serving experience, and mentoring ability.
Sodalis: AI operating system for specialty pharmacy workflows and automation.
5+ YOE5+ years building production ML systems, experience in information extraction/NLP/LLMs, eval and data-pipeline ownership, startup 0→1 experience, pragmatic modeling and clinical domain curiosity.
Experienced in biomolecular modeling and ML for drug discovery; ability to fine-tune protein/structure models, build ML systems that integrate with wet-lab workflows, and evaluate model impact on drug development.