299 inference engineer jobs at 129 companies in Pleasanton, CA

2mo
Save
Mark Applied
Hide
ML Inference Engineer
San Francisco, California, United States
OnsiteFull Time
Reactor
Reactor: Building infrastructure for real-time generative world models.
Strong expertise in ML engineering, PyTorch, CUDA, and high-performance inference; experience with diffusion models and low-latency systems.
PyTorch, TensorRT, TransformerEngine, Nsight, ONNX Runtime, CUDA
1mo
Save
Mark Applied
Hide
Inference Optimization Engineer
Los Altos or United States or Canada
$198k-$286k/yr HybridFull Time
Modular
Modular: Unified software infrastructure and programming language for AI development.
5+ YOE5+ years in distributed systems or performance engineering; experience building reusable tooling; strong technical judgment, communication, and leadership; GPU/kernel, inference engine, Kubernetes, and LLM familiarity helpful.
GPU, ASIC, inference engine, Kubernetes, Modular Cloud, LLM architectures
1mo
Save
Mark Applied
Hide
LLM Inference Engineer
Los Altos, California, United States
OnsiteFull Time
Majestic Labs
Majestic Labs: Developing memory-first AI server platforms for data centers.
3+ YOE3+ years building or operating production LLM inference systems; strong Python and C++; experience with vLLM/SGLang/TensorRT-LLM/Fireworks; distributed inference and performance profiling skills.
vLLM, SGLang, TensorRT-LLM, Fireworks, Python, C++, collective communication library (CCL)
2mo
Save
Mark Applied
Hide
INFERENCE ENGINEER
San Francisco, California, United States
OnsiteFull Time
MakerMaker
MakerMaker: Small San Francis-based AI startup focused on autonomous agents and production ML systems.
3+ YOESenior ML systems engineer with 3+ years building production-grade, large-scale serving infrastructure; strong distributed systems; GPU-accelerated inference; fluent Python and systems languages (C++, CUDA, ROCm or Triton).
Python, C++, CUDA, ROCm, Triton
1mo
Save
Mark Applied
Hide
Senior Software Engineer, Inference
Palo Alto, California, United States
$185k-$250k/yr HybridFull Time
Pika
Pika: AI-powered platform for generating and editing professional videos
5+ YOE5+ years engineering experience in inference acceleration, GPU programming (CUDA, NCCL), model deployment, quantization, attention optimization, and parallelism for production-scale AI systems.
CUDA, NCCL
1mo
Save
Mark Applied
Hide
Principal LLM Inference Engineer
Santa Clara, California, United States
$195k-$285k/yr HybridFull Time
d-Matrix: Develops high-performance semiconductor chips for generative AI inference.
10+ YOEBachelor's in CS/EE (or equivalent) with 10+ years experience (Master/PhD with 6+ years preferred); strong Python and C/C++; experience optimizing LLM inference, quantization, batching, GPU kernel programming and contributor-level work on inference frameworks.
Python, C, C++, vLLM, SGLang, TensorRT-LLM, ONNX Runtime, CUDA, Triton, JAX
2d
Save
Mark Applied
Hide
AI Inference Engineer
San Francisco or United States or Toronto or New York City or Montreal
$165k-$330k/yr HybridFull Time
Baseten
Baseten: Scalable infrastructure platform for deploying and serving AI models.
2+ YOEDegree in CS/Engineering/Math,2+ years experience,production programming (Python preferred),familiarity with ML model lifecycle,strong communication and customer-facing skills.
Python, Docker, Whisper, ComfyUI
1w
Save
Mark Applied
Hide
Staff Engineer, Inference Optimizations
San Francisco, California, United States
$191k-$239k/yr RemoteFull Time
DigitalOcean
DigitalOceanNew York Stock Exchange: DOCN: Simplifies cloud infrastructure for developers, startups, and SMBs.
5+ YOE5+ years in high-performance computing or AI infrastructure, deep GPU and low-level optimization expertise, experience with CUDA/Triton/ROCm, distributed GPU parallelization, and system design for inference workloads.
CUDA, ROCm, TensorRT, OpenAI Triton, AITER, FlashAttention
1w
Save
Mark Applied
Hide
Senior Inference Engineer, GPU Kernel Optimization
Santa Clara or Austin or New York City or Seattle
$184k-$288k/yr OnsiteFull Time
NVIDIA
NVIDIANASDAQ: NVDA: Designs GPU-accelerated computing and artificial intelligence hardware.
6+ YOE6+ years industry experience; strong Python and C++; hands-on GPU profiling (CUPTI, NSYS, NCU); experience with LLM inference frameworks and GPU kernel optimization; advanced degree or equivalent experience.
Python, C++, CUPTI, NSYS, NCU, TRT-LLM, SGLang, vLLM, CUDA, CUTLASS, Triton, PTX, SASS, LLVM, MLIR, ptxas, FlashInfer
2w
Save
Mark Applied
Hide
Inference Infrastructure Engineer, Serving
Palo Alto, California, United States
$275k-$475k/yr OnsiteFull Time
Elorian AI
Elorian AI: AI lab building multimodal models for advanced visual reasoning.
3+ YOE3+ years building low-latency, high-throughput inference serving systems; knowledge of quantization, batching, speculative decoding, KV cache; experience with vLLM/TensorRT-LLM/Triton/SGLang; multi-GPU model parallelism; C++/CUDA/Python; autoscaling and GPU cost optimization.
vLLM, TensorRT-LLM, Triton, SGLang, C++, CUDA, Python
1w
Save
Mark Applied
Hide
Senior Inference Engineer, GPU Kernel Optimization
Santa Clara or Austin or New York City or Seattle
$184k-$288k/yr OnsiteFull Time
NVIDIA
NVIDIANASDAQ: NVDA: Designs graphics processing units and artificial intelligence hardware.
6+ YOEMaster's/PhD or equivalent,6+ years industry experience,agentic AI systems experience,strong Python/C++,GPU profiling (CUPTI,NSYS,NCU),LLM inference frameworks,CUDA/CUTLASS/Triton and PTX/SASS familiarity.
Python, C++, CUPTI, NSYS, NCU, TRT-LLM, SGLang, vLLM, CUDA, CUTLASS, Triton, PTX, SASS, LLVM, MLIR, ptxas
3mo
Save
Mark Applied
Hide
Distributed LLM Inference Engineer
San Francisco or Palo Alto
$170k-$247k/yr HybridFull Time
Anyscale
Anyscale: Cloud platform for scaling distributed machine learning applications.
Familiarity with running ML inference at large scale with high throughput and low latency; experience with PyTorch; solid understanding of distributed systems.
PyTorch, Ray, vLLM, TensorRT-LLM
2mo
Save
Mark Applied
Hide
Inference Optimization ML Engineer
Palo Alto, California, United States
OnsiteFull Time
Rhoda AI
Rhoda AI: Developing generalist robotic intelligence for real-world industrial automation.
3+ YOE3+ years in inference optimization, ML systems; strong PyTorch; experience with quantization, pruning, distillation; familiarity with Triton/TensorRT; CUDA knowledge.
PyTorch, JAX, TensorRT, Triton, CUDA, XLA, TorchServe, vLLM
3w
Save
Mark Applied
Hide
Senior GPU Inference Performance Engineer
Santa Clara, California, United States
$144k-$246k/yr HybridFull Time
AMD
AMDNASDAQ: AMD: Designs and manufactures computer processors and graphics technology.
Hands-on GPU performance engineering experience; proficent in GPU profiling, distributed inference, AI serving frameworks, Python and C/C++; bachelor's degree preferred.
ROCm, rocProfiler, Omniperf, Omnitrace, CUDA, Nsight Systems, Nsight Compute, DCGM, vLLM, SGLang, FP8, MXFP4, INT4, AWQ, GPTQ, RDMA, RoCE, NCCL, RCCL, Pensando, ConnectX, Docker, containerd, Kubernetes, Python, C++, C, HIP, perftest, ib_write_bw, tcpdump
1w
Save
Mark Applied
Hide
Senior Inference Reliability Engineer
San Mateo, California, United States
OnsiteFull Time
Parasail
Parasail: Provides scalable cloud infrastructure for AI model inference.
5+ YOE5+ years production engineering experience operating customer-facing systems; strong SRE and production diagnostics skills; Kubernetes, Linux, distributed systems, and software engineering proficiency; ability to lead incident response and build observability.
Kubernetes, Linux, Python, Go, Java, C++, Rust, vLLM, SGLang, Triton, TensorRT-LLM
3w
Save
Mark Applied
Hide
AI Inference Engineer - Speech
Seattle or San Jose
$152k-$332k/yr HybridFull Time
Zoom
ZoomNasdaq: ZM: Provides a cloud-based platform for video, voice, and collaboration.
3+ YOEMaster's in CS/EE or related,3+ years in speech recognition or model inference,deep learning expertise,experience with Python,C/C++,CUDA,TensorRT,PyTorch,TensorFlow and GPU optimization.
Python, shell, C/C++, PyTorch, TensorFlow, CUDA, TensorRT, CUDA Graphs, NVIDIA GPUs, TPU, BrightHire
2w
Save
Mark Applied
Hide
Applied AI Inference Engineer
San Francisco or Sunnyvale
$250k-$300k/yr OnsiteFull Time
Crusoe
Crusoe: Provides energy-efficient cloud infrastructure powered by stranded and renewable energy.
Experience optimizing LLM inference, production-serving and profiling skills, strong software engineering with Python or C++, familiarity with vLLM/SGLang and CUDA, and ability to work with customers to ship production deployments.
vLLM, SGLang, CUDA, Docker, Kubernetes, Python, C++
1mo
Save
Mark Applied
Hide
Staff AI Inference and Acceleration Engineer
San Jose, California, United States
$180k-$275k/yr OnsiteFull Time
Figure
Figure: Develops autonomous humanoid robots for commercial and residential tasks.
8+ YOEMS/PhD or equivalent, 8+ years in hardware acceleration/ML systems, expertise in inference runtimes, quantization and pruning, profiling and benchmarking, model-to-hardware mapping, and strong C++/Python skills.
ONNX, TFLite, TVM, MLIR, TensorRT, Torch, SNPE/QNN, JAX, CUDA, ROCm, C++, Python
2mo
Save
Mark Applied
Hide
Lead ML Inference Engineer, Advertising
San Jose or Austin
$247k-$486k/yr HybridFull Time
Roku
RokuNASDAQ: ROKU: Operates a TV streaming platform and sells streaming hardware.
10+ YOE5+ MgmtLead the design and development of a state-of-the-art inference platform; 10+ years in distributed systems; ML serving; leadership experience.
High-performance languages, ML frameworks, GPU acceleration, HPC, Distributed systems, Inference platforms, Monitoring tooling
1w
Save
Mark Applied
Hide
Tech Lead Manager, Inference
Redwood City, California, United States
OnsiteFull Time
Luma AI
Luma AI: Develops multimodal AI for video generation and creative production.
8+ YOE8+ years in large-scale distributed systems or ML infrastructure; experience operating inference fleets at thousands-of-GPUs scale; technical leadership; Python, PyTorch, Kubernetes; scheduling, queuing, autoscaling, observability, and SLO ownership.
vLLM, SGLang, TensorRT-LLM, Python, PyTorch, Kubernetes, FFmpeg, RDMA, NVLink, NVIDIA, AMD, TPU, Trainium, Ray, Rust, C++, CUDA, HIP