49 gpu inference performance engineer jobs at 30 companies in Santa Clara, CA

1mo
Save
Mark Applied
Hide
Senior GPU Inference Performance Engineer
Santa Clara, California, United States
$144k-$246k/yr HybridFull Time
AMD
AMDNASDAQ: AMD: Designs and manufactures computer processors and graphics technology.
Hands-on GPU performance engineering experience; proficent in GPU profiling, distributed inference, AI serving frameworks, Python and C/C++; bachelor's degree preferred.
ROCm, rocProfiler, Omniperf, Omnitrace, CUDA, Nsight Systems, Nsight Compute, DCGM, vLLM, SGLang, FP8, MXFP4, INT4, AWQ, GPTQ, RDMA, RoCE, NCCL, RCCL, Pensando, ConnectX, Docker, containerd, Kubernetes, Python, C++, C, HIP, perftest, ib_write_bw, tcpdump
3d
Save
Mark Applied
Hide
Inference Performance Engineer
San Francisco, California, United States
HybridFull Time
Adaption
Adaption: Develops efficient AI systems that adapt in real-time.
5+ YOE5+ years in ML systems, inference infrastructure, or performance engineering; model-serving expertise; Python and systems-language proficiency; and GPU performance experience with measurable cost or latency improvements.
vLLM, SGLang, TensorRT-LLM, Python, C++, Rust, CUDA, NCCL
2w
Save
Mark Applied
Hide
Inference Performance Engineer, AI Inference Configuration Optimization
Santa Clara or United States
$124k-$242k/yr HybridFull Time
NVIDIA
NVIDIANASDAQ: NVDA: Designs GPU-accelerated computing and artificial intelligence hardware.
3+ YOEBachelor's, master's, or doctoral degree in a related field or equivalent experience; 3+ years' engineering experience; GPU profiling, Python, C++/CUDA, and AI inference optimization expertise required.
TensorRT-LLM, SGLang, vLLM, Dynamo, Nsight Systems, Nsight Compute, CUPTI, PyTorch profiler, Python, C++, CUDA, FlashInfer, NCCL, NIXL, NVSHMEM, MLPerf Inference, SemiAnalysis InferenceX
2w
Save
Mark Applied
Hide
Inference Performance Engineer, AI Inference Configuration Optimization
Santa Clara or California
$124k-$242k/yr HybridFull Time
NVIDIA
NVIDIANASDAQ: NVDA: Designs graphics processing units and artificial intelligence hardware.
3+ YOEBS, MS, PhD, or equivalent experience; 3+ years engineering experience; AI inference optimization expertise; GPU profiling; Python; C++/CUDA; reproducible benchmarking and strong communication.
TensorRT-LLM, SGLang, vLLM, Dynamo, Nsight Systems, Nsight Compute, CUPTI, PyTorch profiler, Python, C++, CUDA, FlashInfer, NCCL, NIXL, NVSHMEM, MLPerf Inference, SemiAnalysis InferenceX
2mo
Save
Mark Applied
Hide
Inference Optimization Engineer
Los Altos or United States or Canada
$198k-$286k/yr HybridFull Time
Modular
Modular: Unified software infrastructure and programming language for AI development.
5+ YOE5+ years in distributed systems or performance engineering; experience building reusable tooling; strong technical judgment, communication, and leadership; GPU/kernel, inference engine, Kubernetes, and LLM familiarity helpful.
GPU, ASIC, inference engine, Kubernetes, Modular Cloud, LLM architectures
1w
Save
Mark Applied
Hide
Machine Learning Performance Engineer - Offboard Training & Inference
Sunnyvale or Washington, D.C. or San Diego or Fort Walton Beach or Ann Arbor or London or Stuttgart or Munich or Stockholm or Bangalore or Seoul or Tokyo
$215k-$285k/yr OnsiteFull Time
Applied Intuition
Applied Intuition: Developing software and simulation infrastructure for autonomous vehicles.
ML performance engineering experience with distributed training, batch inference, GPU or accelerator optimization, Python, and C++ or another systems language; strong debugging and analytical skills required.
FSDP, DeepSpeed, Megatron, NCCL, NVIDIA Triton Inference Server, TensorRT, ONNX Runtime, Ray, Python, C++, CUDA, Triton, CUTLASS, Nsight Systems, Nsight Compute, PyTorch Profiler, perf, Kubernetes, Slurm, ROS, OpenCV
1mo
Save
Mark Applied
Hide
Staff Engineer, Inference Optimizations
San Francisco, California, United States
$191k-$239k/yr RemoteFull Time
DigitalOcean
DigitalOceanNew York Stock Exchange: DOCN: Simplifies cloud infrastructure for developers, startups, and SMBs.
5+ YOE5+ years in high-performance computing or AI infrastructure, deep GPU and low-level optimization expertise, experience with CUDA/Triton/ROCm, distributed GPU parallelization, and system design for inference workloads.
CUDA, ROCm, TensorRT, OpenAI Triton, AITER, FlashAttention
3mo
Save
Mark Applied
Hide
ML Inference Engineer
San Francisco, California, United States
OnsiteFull Time
Reactor
Reactor: Building infrastructure for real-time generative world models.
Strong expertise in ML engineering, PyTorch, CUDA, and high-performance inference; experience with diffusion models and low-latency systems.
PyTorch, TensorRT, TransformerEngine, Nsight, ONNX Runtime, CUDA
1w
Save
Mark Applied
Hide
Inference Systems Performance Architect
San Jose, California, United States
$245k-$325k/yr OnsiteFull Time
SambaNova Systems
SambaNova Systems: Develops custom AI hardware and software for enterprise computing.
12+ YOERequires 12+ years in performance engineering, distributed-systems analysis, workload generation, simulation, modeling, technical leadership, cross-functional influence, mentoring, and delivering complex ambiguous projects.
LLM, GPU, RDU, Headspace, Gympass+, One Medical, Employee Assistance Program (EAP)
1mo
Save
Mark Applied
Hide
GPU Kernel Engineer
San Francisco, California, United States
$180k-$280k/yr OnsiteFull Time
TypeSafe AI
TypeSafe AI: Building reliable, general frontier AI models for automation.
Deep CUDA/GPU kernel expertise, experience building and optimizing training and inference kernels, LLM training experience, profiling and eliminating performance bottlenecks.
CUDA, CuTe DSL
3mo
Save
Mark Applied
Hide
Founding Engineer - ML Performance
San Francisco, California, United States
$250k-$395k/yr RemoteFull Time
uRun
uRun: Infrastructure cloud for interactive, stateful AI inference.
Hands-on CUDA, GPU optimization, and large-scale model inference experience; strong systems and performance engineering skills.
CUDA, GPU, NCCL, PyTorch, Triton, TensorRT, CUDA kernels
2w
Save
Mark Applied
Hide
Backend Inference Runtime Engineer Graduate (AML Inference) - 2027 Start
San Jose, California, United States
OnsiteFull Time
ByteDance
ByteDance: Developing AI-driven content platforms and mobile applications.
Bachelor's or master's degree in a technical discipline; proficient in C/C++, Python, CUDA, GPU architecture, deep learning operators, inference compilation, performance analysis, and distributed model inference.
C, C++, Python, CUDA, Nsight, Profiler, vLLM, TensorRT-LLM, SGLang
3mo
Save
Mark Applied
Hide
Performance Engineer
Palo Alto, California, United States
OnsiteFull Time
RadixArk
RadixArk: Building scalable open-source infrastructure for AI training and inference.
Strong systems engineering in performance-critical software; GPU/distributed systems; profiling tools; Python and C++; CUDA/Triton/ROCm/XLA familiarity; LLM inference concepts; ability to debug across software, hardware, and infra layers; strong communication.
CUDA, Triton, Pallas, ROCm, XLA, Python, C++
1mo
Save
Mark Applied
Hide
Sr. Inference Optimization Engineer (local / edge runtime)
Santa Clara or Hillsboro or Folsom or Phoenix
$195k-$361k/yr HybridFull Time
Intel
IntelNasdaq: INTC: Designs and manufactures microprocessors and semiconductor components.
8+ YOE8+ years software development; strong C++ and/or Python; experience with LLM inference, profiling and optimizing CPU/GPU performance; Linux and low-level debugging expertise.
C++, Python, llama.cpp, vLLM, ggml, Vulkan, SYCL, oneAPI, CUDA, Metal, SIMD, Linux, GGUF, AWQ, GPTQ
2mo
Save
Mark Applied
Hide
ML Systems & Performance Engineer
San Francisco, California, United States
OnsiteFull Time
Engram
Engram: Developing persistent memory layers for enterprise AI systems.
5+ YOE5+ years building training/inference systems; strong engineering skills; experience with ML frameworks, GPUs, distributed systems; bachelor's degree or equivalent experience.
PyTorch, JAX, GPUs
2mo
Save
Mark Applied
Hide
Member of Technical Staff, Inference
San Francisco, California, United States
OnsiteFull Time
Radical Numerics
Radical Numerics: Building general biological intelligence models for scientific discovery.
Deep expertise in large-model inference, GPU performance engineering, kernel development (CUDA/Triton), Python and PyTorch, distributed systems, and production model deployment.
CUDA, Triton, Python, PyTorch, vLLM, TensorRT-LLM, SGLang, DeepSpeed
2mo
Save
Mark Applied
Hide
ML Engineer - Inference & Model Deployment
Cupertino, California, United States
$250k-$310k/yr OnsiteFull Time
Hiring.Cafe
Hiring.Cafe: An AI-powered job search engine and aggregator.
Experience deploying and optimizing deep learning models in production, multi-GPU inference, profiling/benchmarking model performance, inference optimization techniques, and cloud/distributed systems familiarity.
vLLM, TensorRT, SGLang, GPU
2w
Save
Mark Applied
Hide
AI Performance Modeling Engineer
Burlingame, California, United States
$180k-$225k/yr HybridFull Time
Quadric
Quadric: Designing licensable processor IP for on-device AI inference.
Strong Python and quantitative modeling skills; computer architecture knowledge; technical writing ability; AI inference or performance modeling expertise; BS, MS, PhD, or equivalent practical experience.
Python, C++, CUDA, Triton, gem5, Timeloop, MAESTRO, Accel-Sim
4d
Save
Mark Applied
Hide
Staff Software Engineer, Inference Performance Optimization, GenAI, DeepMind
Mountain View, California, United States
$207k-$300k/yr OnsiteFull Time
Google
GoogleNASDAQ: GOOGL: Provides online search, advertising, cloud computing, and consumer electronics.
8+ YOEBachelor's degree or equivalent experience, 8 years of software development, Python and C++, AI inference optimization expertise, and experience with serving codebases and performance tradeoffs.
Python, C++, vLLM, TensorRT-LLM, SGLang, Dynamo, PyTorch profiler, GPU, TPU
2mo
Save
Mark Applied
Hide
Software Engineer, ML Performance Optimization
Foster City, California, United States
$192k-$257k/yr OnsiteFull Time
Zoox
ZooxNASDAQ: AMZN: Developing autonomous robotaxis for urban ride-hailing services.
4+ YOE4+ years total exp; 2+ years in large-scale model training or inference; PyTorch; GPU-accelerated inference; profiling tools; Python or C++.
PyTorch, TensorRT, NVIDIA Nsight, Python, C++