12 gpu inference performance engineer jobs at 12 companies in Vacaville, CA

3d
Save
Mark Applied
Hide
Inference Performance Engineer
San Francisco, California, United States
HybridFull Time
Adaption
Adaption: Develops efficient AI systems that adapt in real-time.
5+ YOE5+ years in ML systems, inference infrastructure, or performance engineering; model-serving expertise; Python and systems-language proficiency; and GPU performance experience with measurable cost or latency improvements.
vLLM, SGLang, TensorRT-LLM, Python, C++, Rust, CUDA, NCCL
1mo
Save
Mark Applied
Hide
Staff Engineer, Inference Optimizations
San Francisco, California, United States
$191k-$239k/yr RemoteFull Time
DigitalOcean
DigitalOceanNew York Stock Exchange: DOCN: Simplifies cloud infrastructure for developers, startups, and SMBs.
5+ YOE5+ years in high-performance computing or AI infrastructure, deep GPU and low-level optimization expertise, experience with CUDA/Triton/ROCm, distributed GPU parallelization, and system design for inference workloads.
CUDA, ROCm, TensorRT, OpenAI Triton, AITER, FlashAttention
3mo
Save
Mark Applied
Hide
ML Inference Engineer
San Francisco, California, United States
OnsiteFull Time
Reactor
Reactor: Building infrastructure for real-time generative world models.
Strong expertise in ML engineering, PyTorch, CUDA, and high-performance inference; experience with diffusion models and low-latency systems.
PyTorch, TensorRT, TransformerEngine, Nsight, ONNX Runtime, CUDA
1mo
Save
Mark Applied
Hide
GPU Kernel Engineer
San Francisco, California, United States
$180k-$280k/yr OnsiteFull Time
TypeSafe AI
TypeSafe AI: Building reliable, general frontier AI models for automation.
Deep CUDA/GPU kernel expertise, experience building and optimizing training and inference kernels, LLM training experience, profiling and eliminating performance bottlenecks.
CUDA, CuTe DSL
3mo
Save
Mark Applied
Hide
Founding Engineer - ML Performance
San Francisco, California, United States
$250k-$395k/yr RemoteFull Time
uRun
uRun: Infrastructure cloud for interactive, stateful AI inference.
Hands-on CUDA, GPU optimization, and large-scale model inference experience; strong systems and performance engineering skills.
CUDA, GPU, NCCL, PyTorch, Triton, TensorRT, CUDA kernels
1mo
Save
Mark Applied
Hide
Sr. Inference Optimization Engineer (local / edge runtime)
Santa Clara or Hillsboro or Folsom or Phoenix
$195k-$361k/yr HybridFull Time
Intel
IntelNasdaq: INTC: Designs and manufactures microprocessors and semiconductor components.
8+ YOE8+ years software development; strong C++ and/or Python; experience with LLM inference, profiling and optimizing CPU/GPU performance; Linux and low-level debugging expertise.
C++, Python, llama.cpp, vLLM, ggml, Vulkan, SYCL, oneAPI, CUDA, Metal, SIMD, Linux, GGUF, AWQ, GPTQ
2mo
Save
Mark Applied
Hide
ML Systems & Performance Engineer
San Francisco, California, United States
OnsiteFull Time
Engram
Engram: Developing persistent memory layers for enterprise AI systems.
5+ YOE5+ years building training/inference systems; strong engineering skills; experience with ML frameworks, GPUs, distributed systems; bachelor's degree or equivalent experience.
PyTorch, JAX, GPUs
2mo
Save
Mark Applied
Hide
Member of Technical Staff, Inference
San Francisco, California, United States
OnsiteFull Time
Radical Numerics
Radical Numerics: Building general biological intelligence models for scientific discovery.
Deep expertise in large-model inference, GPU performance engineering, kernel development (CUDA/Triton), Python and PyTorch, distributed systems, and production model deployment.
CUDA, Triton, Python, PyTorch, vLLM, TensorRT-LLM, SGLang, DeepSpeed
2mo
Save
Mark Applied
Hide
Member of Technical Staff, Inference
San Francisco, California, United States
$350k-$500k/yr OnsiteFull Time
Mirendil
Mirendil: Developing frontier artificial intelligence models to accelerate scientific research.
Experience owning inference systems, optimizing inference performance on GPU/accelerator hardware, and extending distributed inference frameworks for high-throughput, low-latency serving.
vLLM, SGLang, TensorRT-LLM
2mo
Save
Mark Applied
Hide
Staff+ Software Engineer, Inference Runtime
San Francisco or Seattle or New York City
$405k-$485k/yr HybridFull Time
Anthropic
Anthropic: Developing safe and reliable artificial intelligence systems.
Senior IC with deep systems or ML infrastructure experience, hands-on performance profiling and optimization, accelerator ecosystem expertise (CUDA/TPU/Trainium), strong software engineering and cross-org alignment skills, and a relevant bachelor’s degree or equivalent.
Rust, Python, CUDA, XLA, Triton, NeuronX, AWS Neuron, Kubernetes, CI/CD
1mo
Save
Mark Applied
Hide
Staff Machine Learning Engineer
California or San Francisco or Mountain View
$167k-$251k/yr RemoteFull Time
Unity
UnityNYSE: U: Provides software for creating real-time 3D interactive content.
5+ YOE5+ years in software/ML engineering with on-device or performance-critical systems; production deployment of transformer/diffusion models on-device; experience with inference runtimes, quantization, operator fusion, and GPU/compute APIs; strong Python; communication and mentoring skills.
WebGPU, WebNN, WGSL, Metal, Vulkan, SPIR-V, CUDA, ONNX Runtime Web, ONNX Runtime, Transformers.js, WebLLM, TensorFlow.js, CoreML, TFLite, ExecuTorch, Chrome, Dawn, PIX, Instruments, Snapdragon Profiler, Nsight, RenderDoc, Python, TypeScript, JavaScript, MLIR, TVM, IREE, XLA, C++, Objective-C, Swift
3mo
Save
Mark Applied
Hide
Member of Technical Staff - Distributed Systems
San Francisco, California, United States
$200k-$300k/yr OnsiteFull Time
Sail Research
Sail Research: Infrastructure platform for long-horizon agentic AI workloads.
Strong distributed systems fundamentals (concurrency, networking, databases, performance engineering). Ability to design, test, and reason about correctness in large-scale systems. Bonus: ML inference and GPU/accelerator experience.
vLLM, SGLang