41 inference optimization engineer jobs at 34 companies in Vallejo, CA

3mo
Save
Mark Applied
Hide
Inference Optimization ML Engineer
Palo Alto, California, United States
OnsiteFull Time
Rhoda AI
Rhoda AI: Developing generalist robotic intelligence for real-world industrial automation.
3+ YOE3+ years in inference optimization, ML systems; strong PyTorch; experience with quantization, pruning, distillation; familiarity with Triton/TensorRT; CUDA knowledge.
PyTorch, JAX, TensorRT, Triton, CUDA, XLA, TorchServe, vLLM
2mo
Save
Mark Applied
Hide
Senior Software Engineer, Inference
Palo Alto, California, United States
$185k-$250k/yr HybridFull Time
Pika
Pika: AI-powered platform for generating and editing professional videos
5+ YOE5+ years engineering experience in inference acceleration, GPU programming (CUDA, NCCL), model deployment, quantization, attention optimization, and parallelism for production-scale AI systems.
CUDA, NCCL
1mo
Save
Mark Applied
Hide
Staff Engineer, Inference Optimizations
San Francisco, California, United States
$191k-$239k/yr RemoteFull Time
DigitalOcean
DigitalOceanNew York Stock Exchange: DOCN: Simplifies cloud infrastructure for developers, startups, and SMBs.
5+ YOE5+ years in high-performance computing or AI infrastructure, deep GPU and low-level optimization expertise, experience with CUDA/Triton/ROCm, distributed GPU parallelization, and system design for inference workloads.
CUDA, ROCm, TensorRT, OpenAI Triton, AITER, FlashAttention
1mo
Save
Mark Applied
Hide
Inference Infrastructure Engineer, Serving
Palo Alto, California, United States
$275k-$475k/yr OnsiteFull Time
Elorian AI
Elorian AI: AI lab building multimodal models for advanced visual reasoning.
3+ YOE3+ years building low-latency, high-throughput inference serving systems; knowledge of quantization, batching, speculative decoding, KV cache; experience with vLLM/TensorRT-LLM/Triton/SGLang; multi-GPU model parallelism; C++/CUDA/Python; autoscaling and GPU cost optimization.
vLLM, TensorRT-LLM, Triton, SGLang, C++, CUDA, Python
1mo
Save
Mark Applied
Hide
Senior Machine Learning Engineer, LLM Inference Optimization
Palo Alto or California
$195k-$262k/yr OnsiteFull Time
Nebius
NebiusNasdaq: NBIS: Builds cloud infrastructure and software for artificial intelligence development.
Expert Python and PyTorch skills, hands-on LLM/VLM inference deployment and optimization, knowledge of modern inference stacks, quantitative reasoning about latency/throughput/cost, and strong communication.
Python, PyTorch, vLLM, SGLang, TensorRT-LLM, Triton Inference Server, NVIDIA Dynamo, Ray Serve, KServe, CUDA, FlashInfer, LMCache, Ray
1mo
Save
Mark Applied
Hide
Applied AI Inference Engineer
San Francisco or Sunnyvale
$250k-$300k/yr OnsiteFull Time
Crusoe
Crusoe: Provides energy-efficient cloud infrastructure powered by stranded and renewable energy.
Experience optimizing LLM inference, production-serving and profiling skills, strong software engineering with Python or C++, familiarity with vLLM/SGLang and CUDA, and ability to work with customers to ship production deployments.
vLLM, SGLang, CUDA, Docker, Kubernetes, Python, C++
3w
Save
Mark Applied
Hide
Research Engineer, Infrastructure, Inference
San Francisco, California, United States
$350k-$475k/yr OnsiteFull Time
Thinking Machines
Thinking Machines: Building AI systems to extend human will and judgment.
Bachelor's in CS or equivalent, strong engineering skills, experience with deep learning frameworks and inference serving, ability to optimize distributed GPU systems and contribute production-quality code.
PyTorch, JAX, SGLang, vLLM, Kubernetes, Ray, SLURM, Triton, DeepSpeed, XLA
1mo
Save
Mark Applied
Hide
Senior Software Engineer, AI Infrastructure - LVM Inference & Evaluation
Redwood City, California, United States
$168k-$205k/yr HybridFull Time
Ambient.ai
Ambient.ai: AI-powered physical security platform for proactive threat detection.
4+ YOE4+ years building infrastructure or production AI systems; strong Python; experience with ML infrastructure, LLM/LVM inference, inference optimization, evaluation frameworks, cloud and GPU workloads; BS/MS or equivalent.
Python, vLLM, Triton Inference Server, CUDA, NCCL, PyTorch, TensorRT, ONNX
2mo
Save
Mark Applied
Hide
Software Engineer, ML Performance Optimization
Foster City, California, United States
$192k-$257k/yr OnsiteFull Time
Zoox
ZooxNASDAQ: AMZN: Developing autonomous robotaxis for urban ride-hailing services.
4+ YOE4+ years total exp; 2+ years in large-scale model training or inference; PyTorch; GPU-accelerated inference; profiling tools; Python or C++.
PyTorch, TensorRT, NVIDIA Nsight, Python, C++
2mo
Save
Mark Applied
Hide
Sr. Lead AI Engineer (Inference Optimization, FM hosting, AI Platform)
San Jose or San Francisco or New York City or Cambridge or McLean
$230k-$286k/yr OnsiteFull Time
Capital One
Capital OneNYSE: COF: Provides credit card, banking, and auto loan services.
6+ YOEBachelor's plus 6 years or master's plus 4 years developing AI/ML technologies, and 6 years programming with Python, Go, Scala, or Java. Cloud AI deployment and team leadership are preferred.
AWS Ultraclusters, Hugging Face, VectorDBs, NeMo Guardrails, PyTorch, Python, Go, Scala, Java, AWS, Google Cloud, Azure, C++, C#, Golang
2mo
Save
Mark Applied
Hide
Member of Technical Staff, Inference
San Francisco, California, United States
$350k-$500k/yr OnsiteFull Time
Mirendil
Mirendil: Developing frontier artificial intelligence models to accelerate scientific research.
Experience owning inference systems, optimizing inference performance on GPU/accelerator hardware, and extending distributed inference frameworks for high-throughput, low-latency serving.
vLLM, SGLang, TensorRT-LLM
1mo
Save
Mark Applied
Hide
Member of Technical Staff — Inference Infrastructure
San Francisco, California, United States
OnsiteFull Time
Causal Labs
Causal Labs: Building physics-based causal AI models for predictive weather intelligence.
Experience building/optimizing inference and serving systems, distributed compute and GPU parallelism knowledge, familiarity with PyTorch/JAX, hardware-aware optimization, and strong engineering/debugging skills.
TensorRT, Kubernetes, Ray, Slurm, PyTorch, JAX, vLLM, SGLang, Triton
2mo
Save
Mark Applied
Hide
Staff+ Software Engineer, Inference Runtime
San Francisco or Seattle or New York City
$405k-$485k/yr HybridFull Time
Anthropic
Anthropic: Developing safe and reliable artificial intelligence systems.
Senior IC with deep systems or ML infrastructure experience, hands-on performance profiling and optimization, accelerator ecosystem expertise (CUDA/TPU/Trainium), strong software engineering and cross-org alignment skills, and a relevant bachelor’s degree or equivalent.
Rust, Python, CUDA, XLA, Triton, NeuronX, AWS Neuron, Kubernetes, CI/CD
2mo
Save
Mark Applied
Hide
Sr. Lead AI Engineer (Inference Optimization, FM hosting, AI Platform)
San Jose or San Francisco or New York City or Cambridge or McLean
$230k-$286k/yr OnsiteFull Time
Capital One
Capital OneNYSE: COF: A diversified financial services providing banking and credit products.
6+ YOEBachelor's degree plus 6 years or master's degree plus 4 years developing AI/ML technologies; 6 years programming with Python, Go, Scala, or Java; cloud AI deployment experience preferred.
AWS Ultraclusters, Hugging Face, VectorDBs, Nemo Guardrails, PyTorch, Python, Go, Scala, Java, Google Cloud, Azure, C++, C#, Golang
3mo
Save
Mark Applied
Hide
Machine Learning Engineer, Inference & Serving (Speech LLM) - San Francisco
San Francisco, California, United States
$180k-$270k/yr HybridFull Time
Plaud
Plaud: Develops AI-powered voice recorders and automated transcription software.
Experience building and deploying high-throughput, ultra-low-latency inference for LLMs or speech models; optimize latency/throughput; manage KV cache; understand GPU memory hierarchies; collaborate across ML and backend teams.
vLLM, TensorRT-LLM, SGLang, NVIDIA Triton Inference Server, WebSockets, WebRTC, CUDA, PTQ, FP8, INT8, AWQ, GPTQ, Tensor Parallelism, Kubernetes
3mo
Save
Mark Applied
Hide
Founding Engineer - ML Performance
San Francisco, California, United States
$250k-$395k/yr RemoteFull Time
uRun
uRun: Infrastructure cloud for interactive, stateful AI inference.
Hands-on CUDA, GPU optimization, and large-scale model inference experience; strong systems and performance engineering skills.
CUDA, GPU, NCCL, PyTorch, Triton, TensorRT, CUDA kernels
3w
Save
Mark Applied
Hide
Staff Software Engineer - Data Cloud Applied ML
San Francisco or Seattle or New York City
$189k-$315k/yr OnsiteFull Time
Rippling
Rippling: Unified platform managing workforce HR, IT, and finance operations
8+ YOE8+ years software engineering experience, distributed systems ownership, experience training/deploying LLMs, model inference optimization, backend skills in Python/Go/Java, and cloud-native infrastructure (Kubernetes).
Python, Go, Java, Kubernetes
3w
Save
Mark Applied
Hide
Senior Machine Learning Engineer
Chicago or New York City or San Francisco or Seattle or Sunnyvale
$182k-$202k/yr OnsiteFull Time
Uber
UberNYSE: UBER: A technology platform for transportation, delivery, and freight.
4+ YOE4+ years building ML models; BS in CS/CE or related; experience with PyTorch, causal inference or constrained optimization preferred; product and marketplace experience a plus.
PyTorch
1mo
Save
Mark Applied
Hide
Member of Technical Staff (Research Engineer)
San Francisco, California, United States
$200k-$400k/yr OnsiteFull Time
Anthrogen
Anthrogen: AI-driven platform for designing and validating synthetic proteins.
Production-grade Python and a systems language, experience with distributed training, GPU optimization, high-throughput data pipelines, inference/serving at scale, strong communication and problem ownership.
Python, PyTorch, JAX, CUDA
3w
Save
Mark Applied
Hide
Senior Computer Vision Engineer
San Francisco or United States
$195k-$255k/yr HybridFull Time
Pano AI
Pano AI: Detects wildfires using AI-powered cameras and satellite intelligence.
5+ YOEMS/PhD in CS/EE/Robotics,5+ years CV/ML industry experience,PyTorch,edge model deployment (NVIDIA Jetson),CUDA/TensorRT/ONNX,Python and C++,experience optimizing inference.
PyTorch, ARM64, CUDA, TensorRT, ONNX, NVIDIA Jetson, Python, C++, DINOv2, DINOv3, SAM, Grounding DINO, Florence