121 inference optimization engineer jobs at 60 companies in South San Francisco, CA

2mo
Save
Mark Applied
Hide
Inference Optimization Engineer
Los Altos or United States or Canada
$198k-$286k/yr HybridFull Time
Modular AI
Modular AINASDAQ: QCOM: AI infrastructure helping developers optimize and deploy AI workloads across heterogeneous computing hardware.
5+ YOE5+ years in distributed systems or performance engineering; experience building reusable tooling; strong technical judgment, communication, and leadership; GPU/kernel, inference engine, Kubernetes, and LLM familiarity helpful.
GPU, ASIC, inference engine, Kubernetes, Modular Cloud, LLM architectures
3mo
Save
Mark Applied
Hide
Inference Optimization ML Engineer
Palo Alto, California, United States
OnsiteFull Time
Rhoda AI
Rhoda AI: Private robotics startup building robot foundation models for autonomous industrial tasks in manufacturing, logistics, automotive, and ecommerce.
3+ YOE3+ years in inference optimization, ML systems; strong PyTorch; experience with quantization, pruning, distillation; familiarity with Triton/TensorRT; CUDA knowledge.
PyTorch, JAX, TensorRT, Triton, CUDA, XLA, TorchServe, vLLM
3w
Save
Mark Applied
Hide
Inference Performance Engineer, Agent Driven Inference Optimization
Santa Clara or United States
$124k-$242k/yr HybridFull Time
NVIDIA
NVIDIANASDAQ: NVDA: Computing platform for AI and accelerated graphics.
3+ YOEBS, MS, or PhD in a relevant technical field or equivalent experience; 3+ years of engineering experience; AI inference optimization expertise; GPU profiling; Python and C++/CUDA skills.
TensorRT-LLM, SGLang, vLLM, Dynamo, Nsight Systems, Nsight Compute, CUPTI, PyTorch profiler, Python, C++, CUDA, FlashInfer, NCCL, NIXL, NVSHMEM, Tensor Cores, TMA, MLPerf Inference, SemiAnalysis InferenceX
1mo
Save
Mark Applied
Hide
Sr. Inference Optimization Engineer (local / edge runtime)
Santa Clara or Hillsboro or Folsom or Phoenix
$195k-$361k/yr HybridFull Time
Intel
IntelNASDAQ: INTC: Design and manufacture of semiconductors and computing technology.
8+ YOE8+ years software development; strong C++ and/or Python; experience with LLM inference, profiling and optimizing CPU/GPU performance; Linux and low-level debugging expertise.
C++, Python, llama.cpp, vLLM, ggml, Vulkan, SYCL, oneAPI, CUDA, Metal, SIMD, Linux, GGUF, AWQ, GPTQ
1mo
Save
Mark Applied
Hide
Senior Inference Engineer, GPU Kernel Optimization
Santa Clara or Austin or New York City or Seattle
$184k-$288k/yr OnsiteFull Time
NVIDIA
NVIDIANASDAQ: NVDA: Computing platform for AI and accelerated graphics.
6+ YOE6+ years industry experience; strong Python and C++; hands-on GPU profiling (CUPTI, NSYS, NCU); experience with LLM inference frameworks and GPU kernel optimization; advanced degree or equivalent experience.
Python, C++, CUPTI, NSYS, NCU, TRT-LLM, SGLang, vLLM, CUDA, CUTLASS, Triton, PTX, SASS, LLVM, MLIR, ptxas, FlashInfer
1mo
Save
Mark Applied
Hide
Senior Inference Engineer, GPU Kernel Optimization
Santa Clara or Austin or New York City or Seattle
$184k-$288k/yr OnsiteFull Time
NVIDIA
NVIDIANASDAQ: NVDA: Computing platform for AI and accelerated graphics.
6+ YOEMaster's/PhD or equivalent,6+ years industry experience,agentic AI systems experience,strong Python/C++,GPU profiling (CUPTI,NSYS,NCU),LLM inference frameworks,CUDA/CUTLASS/Triton and PTX/SASS familiarity.
Python, C++, CUPTI, NSYS, NCU, TRT-LLM, SGLang, vLLM, CUDA, CUTLASS, Triton, PTX, SASS, LLVM, MLIR, ptxas
2mo
Save
Mark Applied
Hide
Senior Software Engineer, Inference
Palo Alto, California, United States
$185k-$250k/yr HybridFull Time
Pika
Pika: AI video creation platform serving creators by generating videos from text, images, and other user content.
5+ YOE5+ years engineering experience in inference acceleration, GPU programming (CUDA, NCCL), model deployment, quantization, attention optimization, and parallelism for production-scale AI systems.
CUDA, NCCL
1mo
Save
Mark Applied
Hide
Staff Engineer, Inference Optimizations
San Francisco, California, United States
$191k-$239k/yr RemoteFull Time
DigitalOcean
DigitalOceanNYSE: DOCN: The AI-Native Cloud purpose-built for inference and agentic workloads.
5+ YOE5+ years in high-performance computing or AI infrastructure, deep GPU and low-level optimization expertise, experience with CUDA/Triton/ROCm, distributed GPU parallelization, and system design for inference workloads.
CUDA, ROCm, TensorRT, OpenAI Triton, AITER, FlashAttention
1mo
Save
Mark Applied
Hide
Principal LLM Inference Engineer
Santa Clara, California, United States
$195k-$285k/yr HybridFull Time
d-Matrix
d-Matrix: Private AI infrastructure serving data centers with inference accelerators, networking, and software.
10+ YOEBachelor's in CS/EE (or equivalent) with 10+ years experience (Master/PhD with 6+ years preferred); strong Python and C/C++; experience optimizing LLM inference, quantization, batching, GPU kernel programming and contributor-level work on inference frameworks.
Python, C, C++, vLLM, SGLang, TensorRT-LLM, ONNX Runtime, CUDA, Triton, JAX
2d
Save
Mark Applied
Hide
Senior Inference Engineer, AGI
Sunnyvale or Boston or Seattle or Los Angeles County
$167k-$260k/yr OnsiteFull Time
Amazon
AmazonNASDAQ: AMZN: Multinational technology focused on e-commerce and cloud computing.
3+ YOERequires 3+ years building machine learning models, 2+ years optimizing neural-model inference, production real-time inference experience, GPU optimization expertise, and Java, C++, or Python programming.
Java, C++, Python, R, scikit-learn, Spark MLLib, MxNet, Tensorflow, numpy, scipy, vLLM, PyTorch, TensorRT-LLM, CUTLASS, Triton, CUDA, PTX, FlashAttention, NCCL, NVLink, AWS Neuron, Trainium, NVIDIA GPU
1mo
Save
Mark Applied
Hide
LLM Inference Engineer
San Francisco or United States
RemoteFull Time
NEAR AI
NEAR AI: Confidential AI infrastructure running verifiable workloads for enterprises, governments, and AI applications.
Expert in LLM inference and serving systems, optimizing throughput/latency/cost for open-source LLMs, deep GPU architecture knowledge, and experience with PyTorch, Triton, CUDA and inference engines like vLLM/SGLang/TensorRT.
SGLang, vLLM, TensorRT, PyTorch, Triton, CuTe, CUDA
1mo
Save
Mark Applied
Hide
Inference Infrastructure Engineer, Serving
Palo Alto, California, United States
$275k-$475k/yr OnsiteFull Time
Elorian
Elorian: AI research lab building multimodal visual-reasoning models for machines, robotics teams, engineers, and scientific organizations.
3+ YOE3+ years building low-latency, high-throughput inference serving systems; knowledge of quantization, batching, speculative decoding, KV cache; experience with vLLM/TensorRT-LLM/Triton/SGLang; multi-GPU model parallelism; C++/CUDA/Python; autoscaling and GPU cost optimization.
vLLM, TensorRT-LLM, Triton, SGLang, C++, CUDA, Python
2d
Save
Mark Applied
Hide
AI Inference Engineer
San Jose or Seattle or United States
$177k-$265k/yr OnsiteFull Time
F5
F5NASDAQ: FFIV: Delivering and securing applications across any multi-cloud environment.
Proficiency in Python, C++, Rust, or Golang; experience with inference tools and cloud infrastructure; expertise in GPU and AI hardware optimization; scalable AI serving and monitoring experience.
vLLM, TGI (Text Generation Inference), NVIDIA Triton, NVIDIA GPUs, CUDA, TensorRT, Apple Silicon, CoreML, TPUs, LPUs, Kubernetes, Python, C++, Rust, Golang, Llama.cpp, Ollama, Docker, AWS, GCP, Azure, Speculative Decoding, PagedAttention, Triton kernels, MLOps, SRE, Time to First Token (TTFT), SLAs
1mo
Save
Mark Applied
Hide
Senior Machine Learning Engineer, LLM Inference Optimization
Palo Alto or California
$195k-$262k/yr OnsiteFull Time
Nebius Group
Nebius GroupNASDAQ: NBIS: Building a full-stack AI cloud infrastructure platform.
Expert Python and PyTorch skills, hands-on LLM/VLM inference deployment and optimization, knowledge of modern inference stacks, quantitative reasoning about latency/throughput/cost, and strong communication.
Python, PyTorch, vLLM, SGLang, TensorRT-LLM, Triton Inference Server, NVIDIA Dynamo, Ray Serve, KServe, CUDA, FlashInfer, LMCache, Ray
2mo
Save
Mark Applied
Hide
LLM Algorithmic Optimization Engineer
San Jose, California, United States
$143k-$186k/yr OnsiteFull Time
NIO
NIONYSE: NIO: Global smart electric vehicle and battery technology provider.
Master's or PhD in a relevant technical field; expertise in GPU/NPU optimization, LLM/VLM architectures, Python, C/C++, PyTorch, inference engines, ONNX, and distributed computing.
Large Language Models (LLMs), Open Neural Network Exchange (ONNX), GPU, NPU, Python, PyTorch, C, C++, Linux kernel, hypervisor
1mo
Save
Mark Applied
Hide
AI Inference Engineer - Speech
Seattle or San Jose
$152k-$332k/yr HybridFull Time
Zoom
ZoomNasdaq Global Select Market: ZM: American publicly traded communications platform serving businesses and individuals with AI-assisted video, voice, chat, and phone services.
3+ YOEMaster's in CS/EE or related,3+ years in speech recognition or model inference,deep learning expertise,experience with Python,C/C++,CUDA,TensorRT,PyTorch,TensorFlow and GPU optimization.
Python, shell, C/C++, PyTorch, TensorFlow, CUDA, TensorRT, CUDA Graphs, NVIDIA GPUs, TPU, BrightHire
1mo
Save
Mark Applied
Hide
Applied AI Inference Engineer
San Francisco or Sunnyvale
$250k-$300k/yr OnsiteFull Time
Crusoe
Crusoe: Vertically integrated AI infrastructure and energy.
Experience optimizing LLM inference, production-serving and profiling skills, strong software engineering with Python or C++, familiarity with vLLM/SGLang and CUDA, and ability to work with customers to ship production deployments.
vLLM, SGLang, CUDA, Docker, Kubernetes, Python, C++
2w
Save
Mark Applied
Hide
Senior/Sr. Staff AI Infrastructure Engineer, Inference & Optimization
San Jose or Mountain View
$170k-$351k/yr OnsiteFull Time
DiDi Autonomous Driving
DiDi Autonomous Driving: Chinese autonomous-driving developing Level 4 self-driving technology, robotaxis, and autonomous trucking logistics for mobility-fleet applications.
3+ YOEMaster's degree or higher in a technical field; 3+ years in HPC, AI infrastructure, model optimization, or embedded deployment; C++, Python, CUDA, OpenMP, inference engines, GPU architectures, and system profiling expertise.
C++, Python, CUDA, OpenMP, TensorRT, ONNX Runtime, vLLM, SGLang, TensorRT-LLM, NVIDIA Hopper, NVIDIA Thor, TGI, LightLLM, PyTorch, INT8, FP8, AWQ, LLaMA, Qwen, GPT
3w
Save
Mark Applied
Hide
Research Engineer, Infrastructure, Inference
San Francisco, California, United States
$350k-$475k/yr OnsiteFull Time
Thinking Machines Lab
Thinking Machines Lab: Private AI research and product building customizable multimodal systems for researchers and the wider public.
Bachelor's in CS or equivalent, strong engineering skills, experience with deep learning frameworks and inference serving, ability to optimize distributed GPU systems and contribute production-quality code.
PyTorch, JAX, SGLang, vLLM, Kubernetes, Ray, SLURM, Triton, DeepSpeed, XLA
1w
Save
Mark Applied
Hide
Staff Software Engineer, Inference Performance Optimization, GenAI, DeepMind
Mountain View, California, United States
$207k-$300k/yr OnsiteFull Time
Google
GoogleNASDAQ: GOOG, GOOGL: Global technology specializing in internet-related services and products.
8+ YOEBachelor's degree or equivalent experience, 8 years of software development, Python and C++, AI inference optimization expertise, and experience with serving codebases and performance tradeoffs.
Python, C++, vLLM, TensorRT-LLM, SGLang, Dynamo, PyTorch profiler, GPU, TPU