Waymo: Autonomous driving technology for ride-hailing and logistics.
4+ YOEBS in CS or equivalent; 4+ years backend experience; experience building distributed backend systems; preferred C++, MS CS; experience with low-latency, large-scale distributed systems; ML/optimization infrastructure and production models.
Unconventional: Developing novel computing hardware for efficient AI acceleration.
MS/PhD (or equivalent) in quantitative field, deep practical experience with ML stack and GPU performance optimization, proficiency in profiling and optimizing large ML codebases.
Senior Data Scientist, Cloud Gaming - Prescriptive Analytics and Optimization
Santa Clara or United States
$184k-$288k/yrHybridFull Time
NVIDIANASDAQ: NVDA: Designs GPU-accelerated computing and artificial intelligence hardware.
6+ YOE6+ years experience (or PhD) in quantitative field, strong prescriptive analytics and optimization background, Python and SQL coding, experience with large-scale data platforms and ML/optimization tooling.
Python, SQL, Delta Lake, Apache Spark, Databricks, MLflow, Grafana, Elasticsearch, Google OR-Tools, Kubeflow
Rhoda AI: Developing generalist robotic intelligence for real-world industrial automation.
3+ YOE3+ years in inference optimization, ML systems; strong PyTorch; experience with quantization, pruning, distillation; familiarity with Triton/TensorRT; CUDA knowledge.
Machine Learning Engineer - AI Compiler Optimization
San Jose, California, United States
OnsiteFull Time
ByteDance: Developing AI-driven content platforms and mobile applications.
Proficient with AI compiler frameworks and GPU/NPU compilation optimization; experience with model import/conversion for PyTorch/TensorFlow and performance tuning for recommendation models.
Research Scientist / Engineer – Performance Optimization
Redwood City, California, United States
OnsiteFull Time
Luma AI: Develops multimodal AI for video generation and creative production.
Expert GPU/CPU/accelerator optimization with Triton/CUDA, strong PyTorch and kernel development, profiling tools experience, deep transformer knowledge, and distributed deployment skills.
Software Engineer, ML Infrastructure, Optimization
Mountain View, California, United States
$160k-$241k/yrOnsiteFull Time
Nuro: Builds autonomous driving software and electric delivery robots.
2+ YOE2+ years in ML optimization infrastructure; experience with quantization, pruning, ML compilers and GPU runtimes; proficient in Python, C++, CUDA and deep learning frameworks (PyTorch, JAX, TensorFlow, Keras).
IntelNasdaq: INTC: Designs and manufactures microprocessors and semiconductor components.
8+ YOE8+ years software development; strong C++ and/or Python; experience with LLM inference, profiling and optimizing CPU/GPU performance; Linux and low-level debugging expertise.
CoreWeaveNASDAQ: CRWV: Cloud platform providing GPU-accelerated infrastructure for AI workloads.
5+ YOE5+ years building HPC/GPU software, hands-on CUDA kernel authoring and optimization, C++/Python coding, GPU profiling, and experience delivering performance at scale.
Hiring.Cafe: An AI-powered job search engine and aggregator.
Experience deploying and optimizing deep learning models in production, multi-GPU inference, profiling/benchmarking model performance, inference optimization techniques, and cloud/distributed systems familiarity.
Software Development Manager, LLM Inference Model Enablement, Neuron SDK
Cupertino, California, United States
$213k-$288k/yrOnsiteFull Time
AmazonNASDAQ: AMZN: Global online retail and cloud computing technology provider.
7+ YOE3+ MgmtManage engineering team to onboard and optimize LLMs for inference on Trainium; strong background in LLM architectures, model performance optimization, and inference techniques; experience with PyTorch and Neuron stack.
Gridmatic: AI-powered platform for optimizing energy trading and battery storage.
Advanced degree in EE/power systems, strong power systems modeling and optimization background, experience with power flow, transmission/congestion analysis, large-scale optimization, and Python programming.
GridCARE: Unlocking electrical grid capacity for AI data center power.
1+ YOEMaster's or equivalent in a quantitative field, 1+ years industry or applied research experience (1–5 years typical), production-quality Python or Julia coding, optimization solvers and simulation experience, strong math and software engineering.
AlphabetNASDAQ: GOOGL: Holding providing internet, software, and AI services.
6+ YOEMaster's or PhD in STEM, 6+ years in modeling and optimization, strong power systems knowledge, numerical methods and solver experience, ability to produce production-ready code, collaborative across disciplines.
NebiusNasdaq: NBIS: Builds cloud infrastructure and software for artificial intelligence development.
Expert Python and PyTorch skills, hands-on LLM/VLM inference deployment and optimization, knowledge of modern inference stacks, quantitative reasoning about latency/throughput/cost, and strong communication.
Python, PyTorch, vLLM, SGLang, TensorRT-LLM, Triton Inference Server, NVIDIA Dynamo, Ray Serve, KServe, CUDA, FlashInfer, LMCache, Ray