102 inference optimization engineer jobs at 53 companies in San Francisco, CA
2mo
Save
Mark Applied
Hide
2mo
Inference Optimization Engineer
Los Altos or United States or Canada
$198k-$286k/yrHybridFull Time
Modular: Unified software infrastructure and programming language for AI development.
5+ YOE5+ years in distributed systems or performance engineering; experience building reusable tooling; strong technical judgment, communication, and leadership; GPU/kernel, inference engine, Kubernetes, and LLM familiarity helpful.
Rhoda AI: Developing generalist robotic intelligence for real-world industrial automation.
3+ YOE3+ years in inference optimization, ML systems; strong PyTorch; experience with quantization, pruning, distillation; familiarity with Triton/TensorRT; CUDA knowledge.
Inference Performance Engineer, AI Inference Configuration Optimization
Santa Clara or United States
$124k-$242k/yrHybridFull Time
NVIDIANASDAQ: NVDA: Designs GPU-accelerated computing and artificial intelligence hardware.
3+ YOEBachelor's, master's, or doctoral degree in a related field or equivalent experience; 3+ years' engineering experience; GPU profiling, Python, C++/CUDA, and AI inference optimization expertise required.
IntelNasdaq: INTC: Designs and manufactures microprocessors and semiconductor components.
8+ YOE8+ years software development; strong C++ and/or Python; experience with LLM inference, profiling and optimizing CPU/GPU performance; Linux and low-level debugging expertise.
Pika: AI-powered platform for generating and editing professional videos
5+ YOE5+ years engineering experience in inference acceleration, GPU programming (CUDA, NCCL), model deployment, quantization, attention optimization, and parallelism for production-scale AI systems.
DigitalOceanNew York Stock Exchange: DOCN: Simplifies cloud infrastructure for developers, startups, and SMBs.
5+ YOE5+ years in high-performance computing or AI infrastructure, deep GPU and low-level optimization expertise, experience with CUDA/Triton/ROCm, distributed GPU parallelization, and system design for inference workloads.
d-Matrix: Develops high-performance semiconductor chips for generative AI inference.
10+ YOEBachelor's in CS/EE (or equivalent) with 10+ years experience (Master/PhD with 6+ years preferred); strong Python and C/C++; experience optimizing LLM inference, quantization, batching, GPU kernel programming and contributor-level work on inference frameworks.
Sunnyvale or Boston or Seattle or Los Angeles County
$193k-$262k/yrOnsiteFull Time
AmazonNASDAQ: AMZN: Global online retail and cloud computing technology provider.
5+ YOERequires 5+ years software development, 4+ years systems architecture, a computer science bachelor's degree, 2+ years neural inference optimization, GPU optimization, real-time systems, and technical leadership experience.
Elorian AI: AI lab building multimodal models for advanced visual reasoning.
3+ YOE3+ years building low-latency, high-throughput inference serving systems; knowledge of quantization, batching, speculative decoding, KV cache; experience with vLLM/TensorRT-LLM/Triton/SGLang; multi-GPU model parallelism; C++/CUDA/Python; autoscaling and GPU cost optimization.
NebiusNasdaq: NBIS: Builds cloud infrastructure and software for artificial intelligence development.
Expert Python and PyTorch skills, hands-on LLM/VLM inference deployment and optimization, knowledge of modern inference stacks, quantitative reasoning about latency/throughput/cost, and strong communication.
Python, PyTorch, vLLM, SGLang, TensorRT-LLM, Triton Inference Server, NVIDIA Dynamo, Ray Serve, KServe, CUDA, FlashInfer, LMCache, Ray
ZoomNasdaq: ZM: Provides a cloud-based platform for video, voice, and collaboration.
3+ YOEMaster's in CS/EE or related,3+ years in speech recognition or model inference,deep learning expertise,experience with Python,C/C++,CUDA,TensorRT,PyTorch,TensorFlow and GPU optimization.
Crusoe: Provides energy-efficient cloud infrastructure powered by stranded and renewable energy.
Experience optimizing LLM inference, production-serving and profiling skills, strong software engineering with Python or C++, familiarity with vLLM/SGLang and CUDA, and ability to work with customers to ship production deployments.
vLLM, SGLang, CUDA, Docker, Kubernetes, Python, C++
3+ YOEMaster's degree or higher in a technical field; 3+ years in HPC, AI infrastructure, model optimization, or embedded deployment; C++, Python, CUDA, OpenMP, inference engines, GPU architectures, and system profiling expertise.
Thinking Machines: Building AI systems to extend human will and judgment.
Bachelor's in CS or equivalent, strong engineering skills, experience with deep learning frameworks and inference serving, ability to optimize distributed GPU systems and contribute production-quality code.
8+ YOEBachelor's degree or equivalent experience, 8 years of software development, Python and C++, AI inference optimization expertise, and experience with serving codebases and performance tradeoffs.
Research Engineer - LLM/VLM Inference Optimization (Seed Infra)
San Jose, California, United States
OnsiteFull Time
ByteDance: Developing AI-driven content platforms and mobile applications.
Bachelor's in CS/EE/Software or related; strong C/C++ and Python; experience with PyTorch or TensorFlow; production LLM/VLM inference optimization experience; familiarity with GPU architecture and containerized server debugging.
Machine Learning Performance Engineer - Offboard Training & Inference
Sunnyvale or Washington, D.C. or San Diego or Fort Walton Beach or Ann Arbor or London or Stuttgart or Munich or Stockholm or Bangalore or Seoul or Tokyo
$215k-$285k/yrOnsiteFull Time
Applied Intuition: Developing software and simulation infrastructure for autonomous vehicles.
ML performance engineering experience with distributed training, batch inference, GPU or accelerator optimization, Python, and C++ or another systems language; strong debugging and analytical skills required.
Hiring.Cafe: An AI-powered job search engine and aggregator.
Experience deploying and optimizing deep learning models in production, multi-GPU inference, profiling/benchmarking model performance, inference optimization techniques, and cloud/distributed systems familiarity.
Senior Software Engineer, AI Infrastructure - LVM Inference & Evaluation
Redwood City, California, United States
$168k-$205k/yrHybridFull Time
Ambient.ai: AI-powered physical security platform for proactive threat detection.
4+ YOE4+ years building infrastructure or production AI systems; strong Python; experience with ML infrastructure, LLM/LVM inference, inference optimization, evaluation frameworks, cloud and GPU workloads; BS/MS or equivalent.