Waymo: Autonomous driving technology for ride-hailing and logistics.
4+ YOEBS in CS or equivalent; 4+ years backend experience; experience building distributed backend systems; preferred C++, MS CS; experience with low-latency, large-scale distributed systems; ML/optimization infrastructure and production models.
Unconventional: Developing novel computing hardware for efficient AI acceleration.
MS/PhD (or equivalent) in quantitative field, deep practical experience with ML stack and GPU performance optimization, proficiency in profiling and optimizing large ML codebases.
Senior Data Scientist, Cloud Gaming - Prescriptive Analytics and Optimization
Santa Clara or United States
$184k-$288k/yrHybridFull Time
NVIDIANASDAQ: NVDA: Designs GPU-accelerated computing and artificial intelligence hardware.
6+ YOE6+ years experience (or PhD) in quantitative field, strong prescriptive analytics and optimization background, Python and SQL coding, experience with large-scale data platforms and ML/optimization tooling.
Python, SQL, Delta Lake, Apache Spark, Databricks, MLflow, Grafana, Elasticsearch, Google OR-Tools, Kubeflow
Rhoda AI: Developing generalist robotic intelligence for real-world industrial automation.
3+ YOE3+ years in inference optimization, ML systems; strong PyTorch; experience with quantization, pruning, distillation; familiarity with Triton/TensorRT; CUDA knowledge.
RobloxNYSE: RBLX: Platform for creating and playing user-generated 3D digital experiences.
6+ YOE6+ years in system design; GPU profiling; CUDA/Triton/TensorRT; ML model optimization for LLMs; Bachelor's in CS/CE/Data Science; collaboration skills.
Quadric: Designing licensable processor IP for on-device AI inference.
MS student in CS/CE or related fields; proficient in C/C++/Python; experience in kernel implementation and optimization; experience in performance profiling.
Machine Learning Engineer - AI Compiler Optimization
San Jose, California, United States
OnsiteFull Time
ByteDance: Developing AI-driven content platforms and mobile applications.
Proficient with AI compiler frameworks and GPU/NPU compilation optimization; experience with model import/conversion for PyTorch/TensorFlow and performance tuning for recommendation models.
Research Scientist / Engineer – Performance Optimization
Redwood City, California, United States
OnsiteFull Time
Luma AI: Develops multimodal AI for video generation and creative production.
Expert GPU/CPU/accelerator optimization with Triton/CUDA, strong PyTorch and kernel development, profiling tools experience, deep transformer knowledge, and distributed deployment skills.
Software Engineer, ML Infrastructure, Optimization
Mountain View, California, United States
$160k-$241k/yrOnsiteFull Time
Nuro: Builds autonomous driving software and electric delivery robots.
2+ YOE2+ years in ML optimization infrastructure; experience with quantization, pruning, ML compilers and GPU runtimes; proficient in Python, C++, CUDA and deep learning frameworks (PyTorch, JAX, TensorFlow, Keras).
IntelNasdaq: INTC: Designs and manufactures microprocessors and semiconductor components.
8+ YOE8+ years software development; strong C++ and/or Python; experience with LLM inference, profiling and optimizing CPU/GPU performance; Linux and low-level debugging expertise.
CoreWeaveNASDAQ: CRWV: Cloud platform providing GPU-accelerated infrastructure for AI workloads.
5+ YOE5+ years building HPC/GPU software, hands-on CUDA kernel authoring and optimization, C++/Python coding, GPU profiling, and experience delivering performance at scale.
Software Development Manager, LLM Inference Model Enablement, Neuron SDK
Cupertino, California, United States
$213k-$288k/yrOnsiteFull Time
AmazonNASDAQ: AMZN: Global online retail and cloud computing technology provider.
7+ YOE3+ MgmtManage engineering team to onboard and optimize LLMs for inference on Trainium; strong background in LLM architectures, model performance optimization, and inference techniques; experience with PyTorch and Neuron stack.
Hiring.Cafe: An AI-powered job search engine and aggregator.
Experience deploying and optimizing deep learning models in production, multi-GPU inference, profiling/benchmarking model performance, inference optimization techniques, and cloud/distributed systems familiarity.
Gridmatic: AI-powered platform for optimizing energy trading and battery storage.
Advanced degree in EE/power systems, strong power systems modeling and optimization background, experience with power flow, transmission/congestion analysis, large-scale optimization, and Python programming.