32 inference optimization engineer jobs at 27 companies in Napa, CA

1mo
Save
Mark Applied
Hide
Staff Engineer, Inference Optimizations
San Francisco, California, United States
$191k-$239k/yr RemoteFull Time
DigitalOcean
DigitalOceanNew York Stock Exchange: DOCN: Simplifies cloud infrastructure for developers, startups, and SMBs.
5+ YOE5+ years in high-performance computing or AI infrastructure, deep GPU and low-level optimization expertise, experience with CUDA/Triton/ROCm, distributed GPU parallelization, and system design for inference workloads.
CUDA, ROCm, TensorRT, OpenAI Triton, AITER, FlashAttention
1mo
Save
Mark Applied
Hide
Applied AI Inference Engineer
San Francisco or Sunnyvale
$250k-$300k/yr OnsiteFull Time
Crusoe
Crusoe: Provides energy-efficient cloud infrastructure powered by stranded and renewable energy.
Experience optimizing LLM inference, production-serving and profiling skills, strong software engineering with Python or C++, familiarity with vLLM/SGLang and CUDA, and ability to work with customers to ship production deployments.
vLLM, SGLang, CUDA, Docker, Kubernetes, Python, C++
2w
Save
Mark Applied
Hide
Research Engineer, Infrastructure, Inference
San Francisco, California, United States
$350k-$475k/yr OnsiteFull Time
Thinking Machines
Thinking Machines: Building AI systems to extend human will and judgment.
Bachelor's in CS or equivalent, strong engineering skills, experience with deep learning frameworks and inference serving, ability to optimize distributed GPU systems and contribute production-quality code.
PyTorch, JAX, SGLang, vLLM, Kubernetes, Ray, SLURM, Triton, DeepSpeed, XLA
2mo
Save
Mark Applied
Hide
Sr. Lead AI Engineer (Inference Optimization, FM hosting, AI Platform)
San Jose or San Francisco or New York City or Cambridge or McLean
$230k-$286k/yr OnsiteFull Time
Capital One
Capital OneNYSE: COF: Provides credit card, banking, and auto loan services.
6+ YOEBachelor's plus 6 years or master's plus 4 years developing AI/ML technologies, and 6 years programming with Python, Go, Scala, or Java. Cloud AI deployment and team leadership are preferred.
AWS Ultraclusters, Hugging Face, VectorDBs, NeMo Guardrails, PyTorch, Python, Go, Scala, Java, AWS, Google Cloud, Azure, C++, C#, Golang
2mo
Save
Mark Applied
Hide
Member of Technical Staff, Inference
San Francisco, California, United States
$350k-$500k/yr OnsiteFull Time
Mirendil
Mirendil: Developing frontier artificial intelligence models to accelerate scientific research.
Experience owning inference systems, optimizing inference performance on GPU/accelerator hardware, and extending distributed inference frameworks for high-throughput, low-latency serving.
vLLM, SGLang, TensorRT-LLM
1mo
Save
Mark Applied
Hide
Member of Technical Staff — Inference Infrastructure
San Francisco, California, United States
OnsiteFull Time
Causal Labs
Causal Labs: Building physics-based causal AI models for predictive weather intelligence.
Experience building/optimizing inference and serving systems, distributed compute and GPU parallelism knowledge, familiarity with PyTorch/JAX, hardware-aware optimization, and strong engineering/debugging skills.
TensorRT, Kubernetes, Ray, Slurm, PyTorch, JAX, vLLM, SGLang, Triton
2mo
Save
Mark Applied
Hide
Staff+ Software Engineer, Inference Runtime
San Francisco or Seattle or New York City
$405k-$485k/yr HybridFull Time
Anthropic
Anthropic: Developing safe and reliable artificial intelligence systems.
Senior IC with deep systems or ML infrastructure experience, hands-on performance profiling and optimization, accelerator ecosystem expertise (CUDA/TPU/Trainium), strong software engineering and cross-org alignment skills, and a relevant bachelor’s degree or equivalent.
Rust, Python, CUDA, XLA, Triton, NeuronX, AWS Neuron, Kubernetes, CI/CD
2mo
Save
Mark Applied
Hide
Sr. Lead AI Engineer (Inference Optimization, FM hosting, AI Platform)
San Jose or San Francisco or New York City or Cambridge or McLean
$230k-$286k/yr OnsiteFull Time
Capital One
Capital OneNYSE: COF: A diversified financial services providing banking and credit products.
6+ YOEBachelor's degree plus 6 years or master's degree plus 4 years developing AI/ML technologies; 6 years programming with Python, Go, Scala, or Java; cloud AI deployment experience preferred.
AWS Ultraclusters, Hugging Face, VectorDBs, Nemo Guardrails, PyTorch, Python, Go, Scala, Java, Google Cloud, Azure, C++, C#, Golang
3mo
Save
Mark Applied
Hide
Machine Learning Engineer, Inference & Serving (Speech LLM) - San Francisco
San Francisco, California, United States
$180k-$270k/yr HybridFull Time
Plaud
Plaud: Develops AI-powered voice recorders and automated transcription software.
Experience building and deploying high-throughput, ultra-low-latency inference for LLMs or speech models; optimize latency/throughput; manage KV cache; understand GPU memory hierarchies; collaborate across ML and backend teams.
vLLM, TensorRT-LLM, SGLang, NVIDIA Triton Inference Server, WebSockets, WebRTC, CUDA, PTQ, FP8, INT8, AWQ, GPTQ, Tensor Parallelism, Kubernetes
3mo
Save
Mark Applied
Hide
Founding Engineer - ML Performance
San Francisco, California, United States
$250k-$395k/yr RemoteFull Time
uRun
uRun: Infrastructure cloud for interactive, stateful AI inference.
Hands-on CUDA, GPU optimization, and large-scale model inference experience; strong systems and performance engineering skills.
CUDA, GPU, NCCL, PyTorch, Triton, TensorRT, CUDA kernels
3w
Save
Mark Applied
Hide
Senior/Staff AI Engineer
Sacramento, California, United States
RemoteFull Time
DataDirect Networks
DataDirect Networks: High-performance storage and data management for AI and HPC.
Experienced engineer with production AI systems ownership, deep systems-level expertise in inference performance, and ability to optimize compute, memory, storage, and serving architecture.
3w
Save
Mark Applied
Hide
Staff Software Engineer - Data Cloud Applied ML
San Francisco or Seattle or New York City
$189k-$315k/yr OnsiteFull Time
Rippling
Rippling: Unified platform managing workforce HR, IT, and finance operations
8+ YOE8+ years software engineering experience, distributed systems ownership, experience training/deploying LLMs, model inference optimization, backend skills in Python/Go/Java, and cloud-native infrastructure (Kubernetes).
Python, Go, Java, Kubernetes
3w
Save
Mark Applied
Hide
Senior Machine Learning Engineer
Chicago or New York City or San Francisco or Seattle or Sunnyvale
$182k-$202k/yr OnsiteFull Time
Uber
UberNYSE: UBER: A technology platform for transportation, delivery, and freight.
4+ YOE4+ years building ML models; BS in CS/CE or related; experience with PyTorch, causal inference or constrained optimization preferred; product and marketplace experience a plus.
PyTorch
1mo
Save
Mark Applied
Hide
Member of Technical Staff (Research Engineer)
San Francisco, California, United States
$200k-$400k/yr OnsiteFull Time
Anthrogen
Anthrogen: AI-driven platform for designing and validating synthetic proteins.
Production-grade Python and a systems language, experience with distributed training, GPU optimization, high-throughput data pipelines, inference/serving at scale, strong communication and problem ownership.
Python, PyTorch, JAX, CUDA
3w
Save
Mark Applied
Hide
Senior Computer Vision Engineer
San Francisco or United States
$195k-$255k/yr HybridFull Time
Pano AI
Pano AI: Detects wildfires using AI-powered cameras and satellite intelligence.
5+ YOEMS/PhD in CS/EE/Robotics,5+ years CV/ML industry experience,PyTorch,edge model deployment (NVIDIA Jetson),CUDA/TensorRT/ONNX,Python and C++,experience optimizing inference.
PyTorch, ARM64, CUDA, TensorRT, ONNX, NVIDIA Jetson, Python, C++, DINOv2, DINOv3, SAM, Grounding DINO, Florence
3mo
Save
Mark Applied
Hide
Forward Deployed Engineer
San Francisco, California, United States
OnsiteFull Time
Reactor
Reactor: Building infrastructure for real-time generative world models.
3+ YOE3+ years in ML engineering or a technical role working with external teams; production-level Python; PyTorch; ML inference optimization; travel willingness.
Python, PyTorch, CUDA, TensorRT
3w
Save
Mark Applied
Hide
Software Engineer (Agentic Systems)
San Francisco, California, United States
$176k-$209k/yr OnsiteFull Time
Dialpad
Dialpad: AI-powered cloud communication and contact center software platform.
10+ YOE10+ years software engineering with technical leadership, distributed systems and LLM/agent experience; mentoring, inference/optimization, retrieval, safety, and productionization skills required.
LangChain, LangGraph, CrewAI, AWS, Google
3mo
Save
Mark Applied
Hide
Founding Software Engineer, Perception
San Francisco, California, United States
$160k-$220k/yr OnsiteFull Time
Kovari
Kovari: Building general-purpose robots for hospitality and physical industries.
Design and deploy perception policies for deployed robots; multimodal data; edge inference optimization; CUDA kernel optimization; sim-to-real experience.
CUDA, TensorRT, ONNX Runtime, TVM, C++, Python
2mo
Save
Mark Applied
Hide
Senior Staff Machine Learning Engineer, LLM/VLM Model Architecture & Optimization
Mountain View or San Francisco
$298k-$368k/yr OnsiteFull Time
Waymo
Waymo: Autonomous driving technology for ride-hailing and logistics.
7+ YOE7+ years ML experience with large-scale model development (LLM/VLM), expertise in on-device inference and hardware acceleration, deep learning frameworks (PyTorch, JAX), large-scale training, and a master's degree in CS/EE or equivalent experience.
PyTorch, JAX
2mo
Save
Mark Applied
Hide
Senior Lead AI Engineer (GenAI Platform, Agentic Infrastructure)
New York or San Francisco or McLean or Cambridge or San Jose or Plano
$209k-$286k/yr OnsiteFull Time
Capital One
Capital OneNYSE: COF: Financial services offering credit cards, banking, and loans.
4+ YOEBachelor's in CS/AI/EE/CE +6 years or Master's +4 years; 6+ years programming with Python/Go/Scala/Java; experience deploying scalable AI on cloud; LLM, inference, similarity search, VectorDBs, guardrails, model evaluation, and optimization experience; leadership and research literacy.
Python, Go, Scala, Java, C++, C#, Golang, AWS Ultraclusters, Huggingface, VectorDBs, Nemo Guardrails, PyTorch, AWS, Google Cloud, Azure