140 inference engineer jobs at 97 companies in Vallejo, CA

2mo
Save
Mark Applied
Hide
ML Inference Engineer
San Francisco, California, United States
OnsiteFull Time
Reactor
Reactor: Building infrastructure for real-time generative world models.
Strong expertise in ML engineering, PyTorch, CUDA, and high-performance inference; experience with diffusion models and low-latency systems.
PyTorch, TensorRT, TransformerEngine, Nsight, ONNX Runtime, CUDA
2mo
Save
Mark Applied
Hide
INFERENCE ENGINEER
San Francisco, California, United States
OnsiteFull Time
MakerMaker
MakerMaker: Small San Francis-based AI startup focused on autonomous agents and production ML systems.
3+ YOESenior ML systems engineer with 3+ years building production-grade, large-scale serving infrastructure; strong distributed systems; GPU-accelerated inference; fluent Python and systems languages (C++, CUDA, ROCm or Triton).
Python, C++, CUDA, ROCm, Triton
1mo
Save
Mark Applied
Hide
Senior Software Engineer, Inference
Palo Alto, California, United States
$185k-$250k/yr HybridFull Time
Pika
Pika: AI-powered platform for generating and editing professional videos
5+ YOE5+ years engineering experience in inference acceleration, GPU programming (CUDA, NCCL), model deployment, quantization, attention optimization, and parallelism for production-scale AI systems.
CUDA, NCCL
2d
Save
Mark Applied
Hide
AI Inference Engineer
San Francisco or United States or Toronto or New York City or Montreal
$165k-$330k/yr HybridFull Time
Baseten
Baseten: Scalable infrastructure platform for deploying and serving AI models.
2+ YOEDegree in CS/Engineering/Math,2+ years experience,production programming (Python preferred),familiarity with ML model lifecycle,strong communication and customer-facing skills.
Python, Docker, Whisper, ComfyUI
1w
Save
Mark Applied
Hide
Staff Engineer, Inference Optimizations
San Francisco, California, United States
$191k-$239k/yr RemoteFull Time
DigitalOcean
DigitalOceanNew York Stock Exchange: DOCN: Simplifies cloud infrastructure for developers, startups, and SMBs.
5+ YOE5+ years in high-performance computing or AI infrastructure, deep GPU and low-level optimization expertise, experience with CUDA/Triton/ROCm, distributed GPU parallelization, and system design for inference workloads.
CUDA, ROCm, TensorRT, OpenAI Triton, AITER, FlashAttention
2w
Save
Mark Applied
Hide
Inference Infrastructure Engineer, Serving
Palo Alto, California, United States
$275k-$475k/yr OnsiteFull Time
Elorian AI
Elorian AI: AI lab building multimodal models for advanced visual reasoning.
3+ YOE3+ years building low-latency, high-throughput inference serving systems; knowledge of quantization, batching, speculative decoding, KV cache; experience with vLLM/TensorRT-LLM/Triton/SGLang; multi-GPU model parallelism; C++/CUDA/Python; autoscaling and GPU cost optimization.
vLLM, TensorRT-LLM, Triton, SGLang, C++, CUDA, Python
3mo
Save
Mark Applied
Hide
Distributed LLM Inference Engineer
San Francisco or Palo Alto
$170k-$247k/yr HybridFull Time
Anyscale
Anyscale: Cloud platform for scaling distributed machine learning applications.
Familiarity with running ML inference at large scale with high throughput and low latency; experience with PyTorch; solid understanding of distributed systems.
PyTorch, Ray, vLLM, TensorRT-LLM
2mo
Save
Mark Applied
Hide
Inference Optimization ML Engineer
Palo Alto, California, United States
OnsiteFull Time
Rhoda AI
Rhoda AI: Developing generalist robotic intelligence for real-world industrial automation.
3+ YOE3+ years in inference optimization, ML systems; strong PyTorch; experience with quantization, pruning, distillation; familiarity with Triton/TensorRT; CUDA knowledge.
PyTorch, JAX, TensorRT, Triton, CUDA, XLA, TorchServe, vLLM
1w
Save
Mark Applied
Hide
Senior Inference Reliability Engineer
San Mateo, California, United States
OnsiteFull Time
Parasail
Parasail: Provides scalable cloud infrastructure for AI model inference.
5+ YOE5+ years production engineering experience operating customer-facing systems; strong SRE and production diagnostics skills; Kubernetes, Linux, distributed systems, and software engineering proficiency; ability to lead incident response and build observability.
Kubernetes, Linux, Python, Go, Java, C++, Rust, vLLM, SGLang, Triton, TensorRT-LLM
2w
Save
Mark Applied
Hide
Applied AI Inference Engineer
San Francisco or Sunnyvale
$250k-$300k/yr OnsiteFull Time
Crusoe
Crusoe: Provides energy-efficient cloud infrastructure powered by stranded and renewable energy.
Experience optimizing LLM inference, production-serving and profiling skills, strong software engineering with Python or C++, familiarity with vLLM/SGLang and CUDA, and ability to work with customers to ship production deployments.
vLLM, SGLang, CUDA, Docker, Kubernetes, Python, C++
1w
Save
Mark Applied
Hide
Tech Lead Manager, Inference
Redwood City, California, United States
OnsiteFull Time
Luma AI
Luma AI: Develops multimodal AI for video generation and creative production.
8+ YOE8+ years in large-scale distributed systems or ML infrastructure; experience operating inference fleets at thousands-of-GPUs scale; technical leadership; Python, PyTorch, Kubernetes; scheduling, queuing, autoscaling, observability, and SLO ownership.
vLLM, SGLang, TensorRT-LLM, Python, PyTorch, Kubernetes, FFmpeg, RDMA, NVLink, NVIDIA, AMD, TPU, Trainium, Ray, Rust, C++, CUDA, HIP
1d
Save
Mark Applied
Hide
Research Engineer, Infrastructure, Inference
San Francisco, California, United States
$350k-$475k/yr OnsiteFull Time
Thinking Machines
Thinking Machines: Building AI systems to extend human will and judgment.
Bachelor's in CS or equivalent, strong engineering skills, experience with deep learning frameworks and inference serving, ability to optimize distributed GPU systems and contribute production-quality code.
PyTorch, JAX, SGLang, vLLM, Kubernetes, Ray, SLURM, Triton, DeepSpeed, XLA
1mo
Save
Mark Applied
Hide
Machine Learning Engineer - Inference Maintainer & Developer Experience
New York City or San Francisco or United States or Europe
RemoteFull Time
Roboflow
Roboflow: Platform for building and deploying custom computer vision models.
5+ YOE5+ years building and operating production ML systems, strong CV/inference foundation, CI/CD and test infra experience, proficiency with PyTorch/TensorFlow/ONNX/TensorRT/vLLM, and experience with image/video processing tools.
inference, PyTorch, TensorFlow, ONNX, TensorRT, vLLM, OpenCV, DeepStream, Pillow, PyAV, GitHub, CI/CD
1mo
Save
Mark Applied
Hide
Member of Technical Staff, Inference
San Francisco, California, United States
OnsiteFull Time
Radical Numerics
Radical Numerics: Building general biological intelligence models for scientific discovery.
Deep expertise in large-model inference, GPU performance engineering, kernel development (CUDA/Triton), Python and PyTorch, distributed systems, and production model deployment.
CUDA, Triton, Python, PyTorch, vLLM, TensorRT-LLM, SGLang, DeepSpeed
1d
Save
Mark Applied
Hide
Forward Deployed Engineer (Inference & Post-Training) - Mandarin Speaking
Singapore or San Francisco
HybridFull Time
Together AI
Together AI: Cloud platform for training and deploying artificial intelligence models.
5+ YOE5+ years experience with inference systems, open-source LLM deployment, and post-training pipelines; expert with inference engines; strong Python skills; Mandarin and English proficiency.
vLLM, TensorRT-LLM, SGLang, Python, LoRA, SFT, DPO, RLHF, GRPO, FlashAttention, Hyena, FlexGen, RedPajama
1w
Save
Mark Applied
Hide
Senior Machine Learning Engineer, LLM Inference Optimization
Palo Alto or California
$195k-$262k/yr OnsiteFull Time
Nebius
NebiusNasdaq: NBIS: Builds cloud infrastructure and software for artificial intelligence development.
Expert Python and PyTorch skills, hands-on LLM/VLM inference deployment and optimization, knowledge of modern inference stacks, quantitative reasoning about latency/throughput/cost, and strong communication.
Python, PyTorch, vLLM, SGLang, TensorRT-LLM, Triton Inference Server, NVIDIA Dynamo, Ray Serve, KServe, CUDA, FlashInfer, LMCache, Ray
2mo
Save
Mark Applied
Hide
Performance Engineer, Inference Systems
San Francisco or New York City or Seattle
$350k-$850k/yr OnsiteFull Time
Anthropic
Anthropic: Developing safe and reliable artificial intelligence systems.
Hands-on performance engineering with Python, data analysis, and cross-layer investigations; strong communication of quantitative results.
Python, SQL, Pandas
3w
Save
Mark Applied
Hide
Senior Software Engineer, AI Infrastructure - LVM Inference & Evaluation
Redwood City, California, United States
$168k-$205k/yr HybridFull Time
Ambient.ai
Ambient.ai: AI-powered physical security platform for proactive threat detection.
4+ YOE4+ years building infrastructure or production AI systems; strong Python; experience with ML infrastructure, LLM/LVM inference, inference optimization, evaluation frameworks, cloud and GPU workloads; BS/MS or equivalent.
Python, vLLM, Triton Inference Server, CUDA, NCCL, PyTorch, TensorRT, ONNX
1mo
Save
Mark Applied
Hide
Member of Technical Staff - Inference Research
New York City or San Francisco
$150k-$350k/yr OnsiteFull Time
Modal
Modal: Serverless cloud platform for running AI and data workloads
Research-leaning or systems background in LLM inference; experience with kernels, quantization, schedulers, and autoscaling; record of shipping research/systems; able to take research bets end-to-end and work onsite in NYC or San Francisco.
Flash Attention 4, Python
2mo
Save
Mark Applied
Hide
AI Platform Engineer, Training and Inference
San Francisco, California, United States
HybridFull Time
Saviynt
Saviynt: Provides AI-powered identity governance and cloud security platforms.
ML platform or MLOps engineer with production Ray experience; LLM serving, distributed training, Python and PyTorch; MLflow/Flyte; Bachelor's degree in CS/Engineering.
Ray Train, Ray Serve, Ray Core, Ray Data, vLLM, SGLang, NVIDIA Triton, TorchTrainer, DDP, NCCL, PPO, RLlib, Flyte, MLflow, Qdrant, Pgvector, PyTorch, Python