32 inference optimization engineer jobs at 27 companies in Napa, CA
1mo
Save
Mark Applied
Hide
1mo
Staff Engineer, Inference Optimizations
San Francisco, California, United States
$191k-$239k/yrRemoteFull Time
DigitalOceanNew York Stock Exchange: DOCN: Simplifies cloud infrastructure for developers, startups, and SMBs.
5+ YOE5+ years in high-performance computing or AI infrastructure, deep GPU and low-level optimization expertise, experience with CUDA/Triton/ROCm, distributed GPU parallelization, and system design for inference workloads.
Crusoe: Provides energy-efficient cloud infrastructure powered by stranded and renewable energy.
Experience optimizing LLM inference, production-serving and profiling skills, strong software engineering with Python or C++, familiarity with vLLM/SGLang and CUDA, and ability to work with customers to ship production deployments.
vLLM, SGLang, CUDA, Docker, Kubernetes, Python, C++
Thinking Machines: Building AI systems to extend human will and judgment.
Bachelor's in CS or equivalent, strong engineering skills, experience with deep learning frameworks and inference serving, ability to optimize distributed GPU systems and contribute production-quality code.
Sr. Lead AI Engineer (Inference Optimization, FM hosting, AI Platform)
San Jose or San Francisco or New York City or Cambridge or McLean
$230k-$286k/yrOnsiteFull Time
Capital OneNYSE: COF: Provides credit card, banking, and auto loan services.
6+ YOEBachelor's plus 6 years or master's plus 4 years developing AI/ML technologies, and 6 years programming with Python, Go, Scala, or Java. Cloud AI deployment and team leadership are preferred.
Anthropic: Developing safe and reliable artificial intelligence systems.
Senior IC with deep systems or ML infrastructure experience, hands-on performance profiling and optimization, accelerator ecosystem expertise (CUDA/TPU/Trainium), strong software engineering and cross-org alignment skills, and a relevant bachelor’s degree or equivalent.
Sr. Lead AI Engineer (Inference Optimization, FM hosting, AI Platform)
San Jose or San Francisco or New York City or Cambridge or McLean
$230k-$286k/yrOnsiteFull Time
Capital OneNYSE: COF: A diversified financial services providing banking and credit products.
6+ YOEBachelor's degree plus 6 years or master's degree plus 4 years developing AI/ML technologies; 6 years programming with Python, Go, Scala, or Java; cloud AI deployment experience preferred.
Machine Learning Engineer, Inference & Serving (Speech LLM) - San Francisco
San Francisco, California, United States
$180k-$270k/yrHybridFull Time
Plaud: Develops AI-powered voice recorders and automated transcription software.
Experience building and deploying high-throughput, ultra-low-latency inference for LLMs or speech models; optimize latency/throughput; manage KV cache; understand GPU memory hierarchies; collaborate across ML and backend teams.
DataDirect Networks: High-performance storage and data management for AI and HPC.
Experienced engineer with production AI systems ownership, deep systems-level expertise in inference performance, and ability to optimize compute, memory, storage, and serving architecture.
Rippling: Unified platform managing workforce HR, IT, and finance operations
8+ YOE8+ years software engineering experience, distributed systems ownership, experience training/deploying LLMs, model inference optimization, backend skills in Python/Go/Java, and cloud-native infrastructure (Kubernetes).
Chicago or New York City or San Francisco or Seattle or Sunnyvale
$182k-$202k/yrOnsiteFull Time
UberNYSE: UBER: A technology platform for transportation, delivery, and freight.
4+ YOE4+ years building ML models; BS in CS/CE or related; experience with PyTorch, causal inference or constrained optimization preferred; product and marketplace experience a plus.
Anthrogen: AI-driven platform for designing and validating synthetic proteins.
Production-grade Python and a systems language, experience with distributed training, GPU optimization, high-throughput data pipelines, inference/serving at scale, strong communication and problem ownership.
Pano AI: Detects wildfires using AI-powered cameras and satellite intelligence.
5+ YOEMS/PhD in CS/EE/Robotics,5+ years CV/ML industry experience,PyTorch,edge model deployment (NVIDIA Jetson),CUDA/TensorRT/ONNX,Python and C++,experience optimizing inference.
Reactor: Building infrastructure for real-time generative world models.
3+ YOE3+ years in ML engineering or a technical role working with external teams; production-level Python; PyTorch; ML inference optimization; travel willingness.
Dialpad: AI-powered cloud communication and contact center software platform.
10+ YOE10+ years software engineering with technical leadership, distributed systems and LLM/agent experience; mentoring, inference/optimization, retrieval, safety, and productionization skills required.
Senior Staff Machine Learning Engineer, LLM/VLM Model Architecture & Optimization
Mountain View or San Francisco
$298k-$368k/yrOnsiteFull Time
Waymo: Autonomous driving technology for ride-hailing and logistics.
7+ YOE7+ years ML experience with large-scale model development (LLM/VLM), expertise in on-device inference and hardware acceleration, deep learning frameworks (PyTorch, JAX), large-scale training, and a master's degree in CS/EE or equivalent experience.
Senior Lead AI Engineer (GenAI Platform, Agentic Infrastructure)
New York or San Francisco or McLean or Cambridge or San Jose or Plano
$209k-$286k/yrOnsiteFull Time
Capital OneNYSE: COF: Financial services offering credit cards, banking, and loans.
4+ YOEBachelor's in CS/AI/EE/CE +6 years or Master's +4 years; 6+ years programming with Python/Go/Scala/Java; experience deploying scalable AI on cloud; LLM, inference, similarity search, VectorDBs, guardrails, model evaluation, and optimization experience; leadership and research literacy.