66 inference optimization engineer jobs at 45 companies in Benicia, CA
2mo
Save
Mark Applied
Hide
2mo
Inference Optimization Engineer
Los Altos or United States or Canada
$198k-$286k/yrHybridFull Time
Modular: Unified software infrastructure and programming language for AI development.
5+ YOE5+ years in distributed systems or performance engineering; experience building reusable tooling; strong technical judgment, communication, and leadership; GPU/kernel, inference engine, Kubernetes, and LLM familiarity helpful.
Rhoda AI: Developing generalist robotic intelligence for real-world industrial automation.
3+ YOE3+ years in inference optimization, ML systems; strong PyTorch; experience with quantization, pruning, distillation; familiarity with Triton/TensorRT; CUDA knowledge.
Inference Performance Engineer, AI Inference Configuration Optimization
Santa Clara or United States
$124k-$242k/yrHybridFull Time
NVIDIANASDAQ: NVDA: Designs GPU-accelerated computing and artificial intelligence hardware.
3+ YOEBachelor's, master's, or doctoral degree in a related field or equivalent experience; 3+ years' engineering experience; GPU profiling, Python, C++/CUDA, and AI inference optimization expertise required.
IntelNasdaq: INTC: Designs and manufactures microprocessors and semiconductor components.
8+ YOE8+ years software development; strong C++ and/or Python; experience with LLM inference, profiling and optimizing CPU/GPU performance; Linux and low-level debugging expertise.
Pika: AI-powered platform for generating and editing professional videos
5+ YOE5+ years engineering experience in inference acceleration, GPU programming (CUDA, NCCL), model deployment, quantization, attention optimization, and parallelism for production-scale AI systems.
DigitalOceanNew York Stock Exchange: DOCN: Simplifies cloud infrastructure for developers, startups, and SMBs.
5+ YOE5+ years in high-performance computing or AI infrastructure, deep GPU and low-level optimization expertise, experience with CUDA/Triton/ROCm, distributed GPU parallelization, and system design for inference workloads.
d-Matrix: Develops high-performance semiconductor chips for generative AI inference.
10+ YOEBachelor's in CS/EE (or equivalent) with 10+ years experience (Master/PhD with 6+ years preferred); strong Python and C/C++; experience optimizing LLM inference, quantization, batching, GPU kernel programming and contributor-level work on inference frameworks.
Sunnyvale or Boston or Seattle or Los Angeles County
$193k-$262k/yrOnsiteFull Time
AmazonNASDAQ: AMZN: Global online retail and cloud computing technology provider.
5+ YOERequires 5+ years software development, 4+ years systems architecture, a computer science bachelor's degree, 2+ years neural inference optimization, GPU optimization, real-time systems, and technical leadership experience.
Elorian AI: AI lab building multimodal models for advanced visual reasoning.
3+ YOE3+ years building low-latency, high-throughput inference serving systems; knowledge of quantization, batching, speculative decoding, KV cache; experience with vLLM/TensorRT-LLM/Triton/SGLang; multi-GPU model parallelism; C++/CUDA/Python; autoscaling and GPU cost optimization.
NebiusNasdaq: NBIS: Builds cloud infrastructure and software for artificial intelligence development.
Expert Python and PyTorch skills, hands-on LLM/VLM inference deployment and optimization, knowledge of modern inference stacks, quantitative reasoning about latency/throughput/cost, and strong communication.
Python, PyTorch, vLLM, SGLang, TensorRT-LLM, Triton Inference Server, NVIDIA Dynamo, Ray Serve, KServe, CUDA, FlashInfer, LMCache, Ray
Crusoe: Provides energy-efficient cloud infrastructure powered by stranded and renewable energy.
Experience optimizing LLM inference, production-serving and profiling skills, strong software engineering with Python or C++, familiarity with vLLM/SGLang and CUDA, and ability to work with customers to ship production deployments.
vLLM, SGLang, CUDA, Docker, Kubernetes, Python, C++
3+ YOEMaster's degree or higher in a technical field; 3+ years in HPC, AI infrastructure, model optimization, or embedded deployment; C++, Python, CUDA, OpenMP, inference engines, GPU architectures, and system profiling expertise.
Thinking Machines: Building AI systems to extend human will and judgment.
Bachelor's in CS or equivalent, strong engineering skills, experience with deep learning frameworks and inference serving, ability to optimize distributed GPU systems and contribute production-quality code.
8+ YOEBachelor's degree or equivalent experience, 8 years of software development, Python and C++, AI inference optimization expertise, and experience with serving codebases and performance tradeoffs.
Machine Learning Performance Engineer - Offboard Training & Inference
Sunnyvale or Washington, D.C. or San Diego or Fort Walton Beach or Ann Arbor or London or Stuttgart or Munich or Stockholm or Bangalore or Seoul or Tokyo
$215k-$285k/yrOnsiteFull Time
Applied Intuition: Developing software and simulation infrastructure for autonomous vehicles.
ML performance engineering experience with distributed training, batch inference, GPU or accelerator optimization, Python, and C++ or another systems language; strong debugging and analytical skills required.
Senior Software Engineer, AI Infrastructure - LVM Inference & Evaluation
Redwood City, California, United States
$168k-$205k/yrHybridFull Time
Ambient.ai: AI-powered physical security platform for proactive threat detection.
4+ YOE4+ years building infrastructure or production AI systems; strong Python; experience with ML infrastructure, LLM/LVM inference, inference optimization, evaluation frameworks, cloud and GPU workloads; BS/MS or equivalent.
Cerebras SystemsNasdaq: CBRS: Manufactures specialized computer chips designed for AI.
8+ YOE8+ years software engineering experience, strong C++ and Python skills, GPU inference experience, Linux, containers and Kubernetes, benchmarking and production optimization for latency-sensitive services.
Sr. Machine Learning Engineer, Foundation Models Inference - Cloud OS & Inference
Santa Clara, California, United States
OnsiteFull Time
AppleNASDAQ: AAPL: Designs and sells consumer electronics, software, and online services.
Build and optimize inference frameworks, services, and tools for large-scale foundation models across cloud infrastructure, including language, vision, and speech models.
Sr. Lead AI Engineer (Inference Optimization, FM hosting, AI Platform)
San Jose or San Francisco or New York City or Cambridge or McLean
$230k-$286k/yrOnsiteFull Time
Capital OneNYSE: COF: Provides credit card, banking, and auto loan services.
6+ YOEBachelor's plus 6 years or master's plus 4 years developing AI/ML technologies, and 6 years programming with Python, Go, Scala, or Java. Cloud AI deployment and team leadership are preferred.