117 inference engineer jobs at 82 companies in Fairfield, CA

2mo
Save
Mark Applied
Hide
ML Inference Engineer
San Francisco, California, United States
OnsiteFull Time
Reactor
Reactor: Building infrastructure for real-time generative world models.
Strong expertise in ML engineering, PyTorch, CUDA, and high-performance inference; experience with diffusion models and low-latency systems.
PyTorch, TensorRT, TransformerEngine, Nsight, ONNX Runtime, CUDA
2mo
Save
Mark Applied
Hide
INFERENCE ENGINEER
San Francisco, California, United States
OnsiteFull Time
MakerMaker
MakerMaker: Small San Francis-based AI startup focused on autonomous agents and production ML systems.
3+ YOESenior ML systems engineer with 3+ years building production-grade, large-scale serving infrastructure; strong distributed systems; GPU-accelerated inference; fluent Python and systems languages (C++, CUDA, ROCm or Triton).
Python, C++, CUDA, ROCm, Triton
2d
Save
Mark Applied
Hide
AI Inference Engineer
San Francisco or United States or Toronto or New York City or Montreal
$165k-$330k/yr HybridFull Time
Baseten
Baseten: Scalable infrastructure platform for deploying and serving AI models.
2+ YOEDegree in CS/Engineering/Math,2+ years experience,production programming (Python preferred),familiarity with ML model lifecycle,strong communication and customer-facing skills.
Python, Docker, Whisper, ComfyUI
1w
Save
Mark Applied
Hide
Staff Engineer, Inference Optimizations
San Francisco, California, United States
$191k-$239k/yr RemoteFull Time
DigitalOcean
DigitalOceanNew York Stock Exchange: DOCN: Simplifies cloud infrastructure for developers, startups, and SMBs.
5+ YOE5+ years in high-performance computing or AI infrastructure, deep GPU and low-level optimization expertise, experience with CUDA/Triton/ROCm, distributed GPU parallelization, and system design for inference workloads.
CUDA, ROCm, TensorRT, OpenAI Triton, AITER, FlashAttention
3mo
Save
Mark Applied
Hide
Distributed LLM Inference Engineer
San Francisco or Palo Alto
$170k-$247k/yr HybridFull Time
Anyscale
Anyscale: Cloud platform for scaling distributed machine learning applications.
Familiarity with running ML inference at large scale with high throughput and low latency; experience with PyTorch; solid understanding of distributed systems.
PyTorch, Ray, vLLM, TensorRT-LLM
1w
Save
Mark Applied
Hide
Senior Inference Reliability Engineer
San Mateo, California, United States
OnsiteFull Time
Parasail
Parasail: Provides scalable cloud infrastructure for AI model inference.
5+ YOE5+ years production engineering experience operating customer-facing systems; strong SRE and production diagnostics skills; Kubernetes, Linux, distributed systems, and software engineering proficiency; ability to lead incident response and build observability.
Kubernetes, Linux, Python, Go, Java, C++, Rust, vLLM, SGLang, Triton, TensorRT-LLM
2w
Save
Mark Applied
Hide
Applied AI Inference Engineer
San Francisco or Sunnyvale
$250k-$300k/yr OnsiteFull Time
Crusoe
Crusoe: Provides energy-efficient cloud infrastructure powered by stranded and renewable energy.
Experience optimizing LLM inference, production-serving and profiling skills, strong software engineering with Python or C++, familiarity with vLLM/SGLang and CUDA, and ability to work with customers to ship production deployments.
vLLM, SGLang, CUDA, Docker, Kubernetes, Python, C++
1d
Save
Mark Applied
Hide
Research Engineer, Infrastructure, Inference
San Francisco, California, United States
$350k-$475k/yr OnsiteFull Time
Thinking Machines
Thinking Machines: Building AI systems to extend human will and judgment.
Bachelor's in CS or equivalent, strong engineering skills, experience with deep learning frameworks and inference serving, ability to optimize distributed GPU systems and contribute production-quality code.
PyTorch, JAX, SGLang, vLLM, Kubernetes, Ray, SLURM, Triton, DeepSpeed, XLA
1mo
Save
Mark Applied
Hide
Machine Learning Engineer - Inference Maintainer & Developer Experience
New York City or San Francisco or United States or Europe
RemoteFull Time
Roboflow
Roboflow: Platform for building and deploying custom computer vision models.
5+ YOE5+ years building and operating production ML systems, strong CV/inference foundation, CI/CD and test infra experience, proficiency with PyTorch/TensorFlow/ONNX/TensorRT/vLLM, and experience with image/video processing tools.
inference, PyTorch, TensorFlow, ONNX, TensorRT, vLLM, OpenCV, DeepStream, Pillow, PyAV, GitHub, CI/CD
1mo
Save
Mark Applied
Hide
Member of Technical Staff, Inference
San Francisco, California, United States
OnsiteFull Time
Radical Numerics
Radical Numerics: Building general biological intelligence models for scientific discovery.
Deep expertise in large-model inference, GPU performance engineering, kernel development (CUDA/Triton), Python and PyTorch, distributed systems, and production model deployment.
CUDA, Triton, Python, PyTorch, vLLM, TensorRT-LLM, SGLang, DeepSpeed
1d
Save
Mark Applied
Hide
Forward Deployed Engineer (Inference & Post-Training) - Mandarin Speaking
Singapore or San Francisco
HybridFull Time
Together AI
Together AI: Cloud platform for training and deploying artificial intelligence models.
5+ YOE5+ years experience with inference systems, open-source LLM deployment, and post-training pipelines; expert with inference engines; strong Python skills; Mandarin and English proficiency.
vLLM, TensorRT-LLM, SGLang, Python, LoRA, SFT, DPO, RLHF, GRPO, FlashAttention, Hyena, FlexGen, RedPajama
2mo
Save
Mark Applied
Hide
Performance Engineer, Inference Systems
San Francisco or New York City or Seattle
$350k-$850k/yr OnsiteFull Time
Anthropic
Anthropic: Developing safe and reliable artificial intelligence systems.
Hands-on performance engineering with Python, data analysis, and cross-layer investigations; strong communication of quantitative results.
Python, SQL, Pandas
1mo
Save
Mark Applied
Hide
Member of Technical Staff - Inference Research
New York City or San Francisco
$150k-$350k/yr OnsiteFull Time
Modal
Modal: Serverless cloud platform for running AI and data workloads
Research-leaning or systems background in LLM inference; experience with kernels, quantization, schedulers, and autoscaling; record of shipping research/systems; able to take research bets end-to-end and work onsite in NYC or San Francisco.
Flash Attention 4, Python
2mo
Save
Mark Applied
Hide
AI Platform Engineer, Training and Inference
San Francisco, California, United States
HybridFull Time
Saviynt
Saviynt: Provides AI-powered identity governance and cloud security platforms.
ML platform or MLOps engineer with production Ray experience; LLM serving, distributed training, Python and PyTorch; MLflow/Flyte; Bachelor's degree in CS/Engineering.
Ray Train, Ray Serve, Ray Core, Ray Data, vLLM, SGLang, NVIDIA Triton, TorchTrainer, DDP, NCCL, PPO, RLlib, Flyte, MLflow, Qdrant, Pgvector, PyTorch, Python
2mo
Save
Mark Applied
Hide
Engineering Manager, Model Inference
San Francisco or New York or Pittsburgh
$220k-$270k/yr HybridFull Time
Abridge
Abridge: Automates medical documentation through AI-powered speech analysis
5+ YOE1+ Mgmt5+ years engineering with 1+ year in technical leadership; ML systems and inference experience; GPU, latency, throughput expertise; strong people leadership and collaboration.
PyTorch, TensorRT, vLLM, TensorFlow, GPU
1mo
Save
Mark Applied
Hide
Member of Technical Staff, Inference
San Francisco, California, United States
$350k-$500k/yr OnsiteFull Time
Mirendil
Mirendil: Developing frontier artificial intelligence models to accelerate scientific research.
Experience owning inference systems, optimizing inference performance on GPU/accelerator hardware, and extending distributed inference frameworks for high-throughput, low-latency serving.
vLLM, SGLang, TensorRT-LLM
1w
Save
Mark Applied
Hide
Member of Technical Staff (Inference)
San Francisco, California, United States
OnsiteFull Time
Artificial Analysis
Artificial Analysis: Independent AI benchmarking and performance analysis platform.
3+ YOE3+ years professional experience (min 2 years with inference providers/neoclouds), proficiency in Python and data analysis, hands-on familiarity with vLLM,SGLang,TensorRT-LLM and inference APIs, strong analytical skills, and fluency with inference performance metrics.
Python, vLLM, SGLang, TensorRT-LLM
1mo
Save
Mark Applied
Hide
Machine Learning Engineer, Causal Inference, Level 5
Santa Monica or Los Angeles or San Francisco or Palo Alto or New York City
$178k-$313k/yr OnsiteFull Time
Snap
SnapNYSE: SNAP: Develops social media applications and augmented reality technology.
5+ YOE5+ years post-Bachelor's ML experience with expertise in causal inference, experimentation, uplift modeling, and productionizing models; proficient in Python, pandas, NumPy, scikit-learn; strong communication and mentorship skills.
Python, pandas, NumPy, scikit-learn, CausalM, CausalML, EconML, DoWhy
1mo
Save
Mark Applied
Hide
Machine Learning Engineer, Causal Inference, Level 5
Los Angeles or San Francisco or Palo Alto or New York or Santa Monica
$178k-$313k/yr OnsiteFull Time
Snap
SnapNYSE: SNAP: Provides visual messaging software and augmented reality wearable devices.
5+ YOE5+ years post-Bachelor's ML experience (or Master's/PhD with reduced experience), strong causal inference and experimentation experience, proficiency in Python and ML libraries, production ML experience, strong communication and mentorship skills.
Python, pandas, NumPy, scikit-learn, CausalM, CausalML, EconML, DoWhy
2mo
Save
Mark Applied
Hide
Member of Technical Staff - Inference
San Francisco, California, United States
$200k-$300k/yr OnsiteFull Time
Sail Research
Sail Research: Infrastructure platform for long-horizon agentic AI workloads.
Deep knowledge of LLM mechanics and MLSys research, experience with inference engines and GPU profiling, familiarity with tile-based GPU programming (Triton/CUTLASS/ThunderKittens), strong communication and critical thinking.
vLLM, SGLang, NSys, Triton, CUTLASS, ThunderKittens