121 inference optimization engineer jobs at 60 companies in South San Francisco, CA
2mo
Save
Mark Applied
Hide
2mo
Inference Optimization Engineer
Los Altos or United States or Canada
$198k-$286k/yrHybridFull Time
Modular AINASDAQ: QCOM: AI infrastructure helping developers optimize and deploy AI workloads across heterogeneous computing hardware.
5+ YOE5+ years in distributed systems or performance engineering; experience building reusable tooling; strong technical judgment, communication, and leadership; GPU/kernel, inference engine, Kubernetes, and LLM familiarity helpful.
Rhoda AI: Private robotics startup building robot foundation models for autonomous industrial tasks in manufacturing, logistics, automotive, and ecommerce.
3+ YOE3+ years in inference optimization, ML systems; strong PyTorch; experience with quantization, pruning, distillation; familiarity with Triton/TensorRT; CUDA knowledge.
NVIDIANASDAQ: NVDA: Computing platform for AI and accelerated graphics.
3+ YOEBS, MS, or PhD in a relevant technical field or equivalent experience; 3+ years of engineering experience; AI inference optimization expertise; GPU profiling; Python and C++/CUDA skills.
IntelNASDAQ: INTC: Design and manufacture of semiconductors and computing technology.
8+ YOE8+ years software development; strong C++ and/or Python; experience with LLM inference, profiling and optimizing CPU/GPU performance; Linux and low-level debugging expertise.
NVIDIANASDAQ: NVDA: Computing platform for AI and accelerated graphics.
6+ YOE6+ years industry experience; strong Python and C++; hands-on GPU profiling (CUPTI, NSYS, NCU); experience with LLM inference frameworks and GPU kernel optimization; advanced degree or equivalent experience.
NVIDIANASDAQ: NVDA: Computing platform for AI and accelerated graphics.
6+ YOEMaster's/PhD or equivalent,6+ years industry experience,agentic AI systems experience,strong Python/C++,GPU profiling (CUPTI,NSYS,NCU),LLM inference frameworks,CUDA/CUTLASS/Triton and PTX/SASS familiarity.
Pika: AI video creation platform serving creators by generating videos from text, images, and other user content.
5+ YOE5+ years engineering experience in inference acceleration, GPU programming (CUDA, NCCL), model deployment, quantization, attention optimization, and parallelism for production-scale AI systems.
DigitalOceanNYSE: DOCN: The AI-Native Cloud purpose-built for inference and agentic workloads.
5+ YOE5+ years in high-performance computing or AI infrastructure, deep GPU and low-level optimization expertise, experience with CUDA/Triton/ROCm, distributed GPU parallelization, and system design for inference workloads.
d-Matrix: Private AI infrastructure serving data centers with inference accelerators, networking, and software.
10+ YOEBachelor's in CS/EE (or equivalent) with 10+ years experience (Master/PhD with 6+ years preferred); strong Python and C/C++; experience optimizing LLM inference, quantization, batching, GPU kernel programming and contributor-level work on inference frameworks.
Sunnyvale or Boston or Seattle or Los Angeles County
$167k-$260k/yrOnsiteFull Time
AmazonNASDAQ: AMZN: Multinational technology focused on e-commerce and cloud computing.
3+ YOERequires 3+ years building machine learning models, 2+ years optimizing neural-model inference, production real-time inference experience, GPU optimization expertise, and Java, C++, or Python programming.
NEAR AI: Confidential AI infrastructure running verifiable workloads for enterprises, governments, and AI applications.
Expert in LLM inference and serving systems, optimizing throughput/latency/cost for open-source LLMs, deep GPU architecture knowledge, and experience with PyTorch, Triton, CUDA and inference engines like vLLM/SGLang/TensorRT.
SGLang, vLLM, TensorRT, PyTorch, Triton, CuTe, CUDA
Elorian: AI research lab building multimodal visual-reasoning models for machines, robotics teams, engineers, and scientific organizations.
3+ YOE3+ years building low-latency, high-throughput inference serving systems; knowledge of quantization, batching, speculative decoding, KV cache; experience with vLLM/TensorRT-LLM/Triton/SGLang; multi-GPU model parallelism; C++/CUDA/Python; autoscaling and GPU cost optimization.
F5NASDAQ: FFIV: Delivering and securing applications across any multi-cloud environment.
Proficiency in Python, C++, Rust, or Golang; experience with inference tools and cloud infrastructure; expertise in GPU and AI hardware optimization; scalable AI serving and monitoring experience.
vLLM, TGI (Text Generation Inference), NVIDIA Triton, NVIDIA GPUs, CUDA, TensorRT, Apple Silicon, CoreML, TPUs, LPUs, Kubernetes, Python, C++, Rust, Golang, Llama.cpp, Ollama, Docker, AWS, GCP, Azure, Speculative Decoding, PagedAttention, Triton kernels, MLOps, SRE, Time to First Token (TTFT), SLAs
Nebius GroupNASDAQ: NBIS: Building a full-stack AI cloud infrastructure platform.
Expert Python and PyTorch skills, hands-on LLM/VLM inference deployment and optimization, knowledge of modern inference stacks, quantitative reasoning about latency/throughput/cost, and strong communication.
Python, PyTorch, vLLM, SGLang, TensorRT-LLM, Triton Inference Server, NVIDIA Dynamo, Ray Serve, KServe, CUDA, FlashInfer, LMCache, Ray
NIONYSE: NIO: Global smart electric vehicle and battery technology provider.
Master's or PhD in a relevant technical field; expertise in GPU/NPU optimization, LLM/VLM architectures, Python, C/C++, PyTorch, inference engines, ONNX, and distributed computing.
Large Language Models (LLMs), Open Neural Network Exchange (ONNX), GPU, NPU, Python, PyTorch, C, C++, Linux kernel, hypervisor
ZoomNasdaq Global Select Market: ZM: American publicly traded communications platform serving businesses and individuals with AI-assisted video, voice, chat, and phone services.
3+ YOEMaster's in CS/EE or related,3+ years in speech recognition or model inference,deep learning expertise,experience with Python,C/C++,CUDA,TensorRT,PyTorch,TensorFlow and GPU optimization.
Crusoe: Vertically integrated AI infrastructure and energy.
Experience optimizing LLM inference, production-serving and profiling skills, strong software engineering with Python or C++, familiarity with vLLM/SGLang and CUDA, and ability to work with customers to ship production deployments.
vLLM, SGLang, CUDA, Docker, Kubernetes, Python, C++
Senior/Sr. Staff AI Infrastructure Engineer, Inference & Optimization
San Jose or Mountain View
$170k-$351k/yrOnsiteFull Time
DiDi Autonomous Driving: Chinese autonomous-driving developing Level 4 self-driving technology, robotaxis, and autonomous trucking logistics for mobility-fleet applications.
3+ YOEMaster's degree or higher in a technical field; 3+ years in HPC, AI infrastructure, model optimization, or embedded deployment; C++, Python, CUDA, OpenMP, inference engines, GPU architectures, and system profiling expertise.
Thinking Machines Lab: Private AI research and product building customizable multimodal systems for researchers and the wider public.
Bachelor's in CS or equivalent, strong engineering skills, experience with deep learning frameworks and inference serving, ability to optimize distributed GPU systems and contribute production-quality code.
GoogleNASDAQ: GOOG, GOOGL: Global technology specializing in internet-related services and products.
8+ YOEBachelor's degree or equivalent experience, 8 years of software development, Python and C++, AI inference optimization expertise, and experience with serving codebases and performance tradeoffs.