202 inference engineer jobs at 66 companies in Tracy, CA

1mo
Save
Mark Applied
Hide
Inference Optimization Engineer
Los Altos or United States or Canada
$198k-$286k/yr HybridFull Time
Modular
Modular: Unified software infrastructure and programming language for AI development.
5+ YOE5+ years in distributed systems or performance engineering; experience building reusable tooling; strong technical judgment, communication, and leadership; GPU/kernel, inference engine, Kubernetes, and LLM familiarity helpful.
GPU, ASIC, inference engine, Kubernetes, Modular Cloud, LLM architectures
1mo
Save
Mark Applied
Hide
LLM Inference Engineer
Los Altos, California, United States
OnsiteFull Time
Majestic Labs
Majestic Labs: Developing memory-first AI server platforms for data centers.
3+ YOE3+ years building or operating production LLM inference systems; strong Python and C++; experience with vLLM/SGLang/TensorRT-LLM/Fireworks; distributed inference and performance profiling skills.
vLLM, SGLang, TensorRT-LLM, Fireworks, Python, C++, collective communication library (CCL)
1mo
Save
Mark Applied
Hide
Senior Software Engineer, Inference
Palo Alto, California, United States
$185k-$250k/yr HybridFull Time
Pika
Pika: AI-powered platform for generating and editing professional videos
5+ YOE5+ years engineering experience in inference acceleration, GPU programming (CUDA, NCCL), model deployment, quantization, attention optimization, and parallelism for production-scale AI systems.
CUDA, NCCL
1mo
Save
Mark Applied
Hide
Principal LLM Inference Engineer
Santa Clara, California, United States
$195k-$285k/yr HybridFull Time
d-Matrix: Develops high-performance semiconductor chips for generative AI inference.
10+ YOEBachelor's in CS/EE (or equivalent) with 10+ years experience (Master/PhD with 6+ years preferred); strong Python and C/C++; experience optimizing LLM inference, quantization, batching, GPU kernel programming and contributor-level work on inference frameworks.
Python, C, C++, vLLM, SGLang, TensorRT-LLM, ONNX Runtime, CUDA, Triton, JAX
1w
Save
Mark Applied
Hide
Senior Inference Engineer, GPU Kernel Optimization
Santa Clara or Austin or New York City or Seattle
$184k-$288k/yr OnsiteFull Time
NVIDIA
NVIDIANASDAQ: NVDA: Designs GPU-accelerated computing and artificial intelligence hardware.
6+ YOE6+ years industry experience; strong Python and C++; hands-on GPU profiling (CUPTI, NSYS, NCU); experience with LLM inference frameworks and GPU kernel optimization; advanced degree or equivalent experience.
Python, C++, CUPTI, NSYS, NCU, TRT-LLM, SGLang, vLLM, CUDA, CUTLASS, Triton, PTX, SASS, LLVM, MLIR, ptxas, FlashInfer
2w
Save
Mark Applied
Hide
Inference Infrastructure Engineer, Serving
Palo Alto, California, United States
$275k-$475k/yr OnsiteFull Time
Elorian AI
Elorian AI: AI lab building multimodal models for advanced visual reasoning.
3+ YOE3+ years building low-latency, high-throughput inference serving systems; knowledge of quantization, batching, speculative decoding, KV cache; experience with vLLM/TensorRT-LLM/Triton/SGLang; multi-GPU model parallelism; C++/CUDA/Python; autoscaling and GPU cost optimization.
vLLM, TensorRT-LLM, Triton, SGLang, C++, CUDA, Python
1w
Save
Mark Applied
Hide
Senior Inference Engineer, GPU Kernel Optimization
Santa Clara or Austin or New York City or Seattle
$184k-$288k/yr OnsiteFull Time
NVIDIA
NVIDIANASDAQ: NVDA: Designs graphics processing units and artificial intelligence hardware.
6+ YOEMaster's/PhD or equivalent,6+ years industry experience,agentic AI systems experience,strong Python/C++,GPU profiling (CUPTI,NSYS,NCU),LLM inference frameworks,CUDA/CUTLASS/Triton and PTX/SASS familiarity.
Python, C++, CUPTI, NSYS, NCU, TRT-LLM, SGLang, vLLM, CUDA, CUTLASS, Triton, PTX, SASS, LLVM, MLIR, ptxas
3mo
Save
Mark Applied
Hide
Distributed LLM Inference Engineer
San Francisco or Palo Alto
$170k-$247k/yr HybridFull Time
Anyscale
Anyscale: Cloud platform for scaling distributed machine learning applications.
Familiarity with running ML inference at large scale with high throughput and low latency; experience with PyTorch; solid understanding of distributed systems.
PyTorch, Ray, vLLM, TensorRT-LLM
2mo
Save
Mark Applied
Hide
Inference Optimization ML Engineer
Palo Alto, California, United States
OnsiteFull Time
Rhoda AI
Rhoda AI: Developing generalist robotic intelligence for real-world industrial automation.
3+ YOE3+ years in inference optimization, ML systems; strong PyTorch; experience with quantization, pruning, distillation; familiarity with Triton/TensorRT; CUDA knowledge.
PyTorch, JAX, TensorRT, Triton, CUDA, XLA, TorchServe, vLLM
3w
Save
Mark Applied
Hide
Senior GPU Inference Performance Engineer
Santa Clara, California, United States
$144k-$246k/yr HybridFull Time
AMD
AMDNASDAQ: AMD: Designs and manufactures computer processors and graphics technology.
Hands-on GPU performance engineering experience; proficent in GPU profiling, distributed inference, AI serving frameworks, Python and C/C++; bachelor's degree preferred.
ROCm, rocProfiler, Omniperf, Omnitrace, CUDA, Nsight Systems, Nsight Compute, DCGM, vLLM, SGLang, FP8, MXFP4, INT4, AWQ, GPTQ, RDMA, RoCE, NCCL, RCCL, Pensando, ConnectX, Docker, containerd, Kubernetes, Python, C++, C, HIP, perftest, ib_write_bw, tcpdump
3w
Save
Mark Applied
Hide
AI Inference Engineer - Speech
Seattle or San Jose
$152k-$332k/yr HybridFull Time
Zoom
ZoomNasdaq: ZM: Provides a cloud-based platform for video, voice, and collaboration.
3+ YOEMaster's in CS/EE or related,3+ years in speech recognition or model inference,deep learning expertise,experience with Python,C/C++,CUDA,TensorRT,PyTorch,TensorFlow and GPU optimization.
Python, shell, C/C++, PyTorch, TensorFlow, CUDA, TensorRT, CUDA Graphs, NVIDIA GPUs, TPU, BrightHire
2w
Save
Mark Applied
Hide
Applied AI Inference Engineer
San Francisco or Sunnyvale
$250k-$300k/yr OnsiteFull Time
Crusoe
Crusoe: Provides energy-efficient cloud infrastructure powered by stranded and renewable energy.
Experience optimizing LLM inference, production-serving and profiling skills, strong software engineering with Python or C++, familiarity with vLLM/SGLang and CUDA, and ability to work with customers to ship production deployments.
vLLM, SGLang, CUDA, Docker, Kubernetes, Python, C++
1mo
Save
Mark Applied
Hide
Staff AI Inference and Acceleration Engineer
San Jose, California, United States
$180k-$275k/yr OnsiteFull Time
Figure
Figure: Develops autonomous humanoid robots for commercial and residential tasks.
8+ YOEMS/PhD or equivalent, 8+ years in hardware acceleration/ML systems, expertise in inference runtimes, quantization and pruning, profiling and benchmarking, model-to-hardware mapping, and strong C++/Python skills.
ONNX, TFLite, TVM, MLIR, TensorRT, Torch, SNPE/QNN, JAX, CUDA, ROCm, C++, Python
2mo
Save
Mark Applied
Hide
Lead ML Inference Engineer, Advertising
San Jose or Austin
$247k-$486k/yr HybridFull Time
Roku
RokuNASDAQ: ROKU: Operates a TV streaming platform and sells streaming hardware.
10+ YOE5+ MgmtLead the design and development of a state-of-the-art inference platform; 10+ years in distributed systems; ML serving; leadership experience.
High-performance languages, ML frameworks, GPU acceleration, HPC, Distributed systems, Inference platforms, Monitoring tooling
1w
Save
Mark Applied
Hide
Tech Lead Manager, Inference
Redwood City, California, United States
OnsiteFull Time
Luma AI
Luma AI: Develops multimodal AI for video generation and creative production.
8+ YOE8+ years in large-scale distributed systems or ML infrastructure; experience operating inference fleets at thousands-of-GPUs scale; technical leadership; Python, PyTorch, Kubernetes; scheduling, queuing, autoscaling, observability, and SLO ownership.
vLLM, SGLang, TensorRT-LLM, Python, PyTorch, Kubernetes, FFmpeg, RDMA, NVLink, NVIDIA, AMD, TPU, Trainium, Ray, Rust, C++, CUDA, HIP
1w
Save
Mark Applied
Hide
Sr. Inference Optimization Engineer (local / edge runtime)
Santa Clara or Hillsboro or Folsom or Phoenix
$195k-$361k/yr HybridFull Time
Intel
IntelNasdaq: INTC: Designs and manufactures microprocessors and semiconductor components.
8+ YOE8+ years software development; strong C++ and/or Python; experience with LLM inference, profiling and optimizing CPU/GPU performance; Linux and low-level debugging expertise.
C++, Python, llama.cpp, vLLM, ggml, Vulkan, SYCL, oneAPI, CUDA, Metal, SIMD, Linux, GGUF, AWQ, GPTQ
3w
Save
Mark Applied
Hide
AI/ML Engineer - Model Inference
Sunnyvale, California, United States
$118k-$221k/yr HybridFull Time
General Motors
General MotorsNYSE: GM: Manufactures and sells automobiles and automotive parts globally.
Experience building production-scale ML data pipelines, featurization, embedding and inference systems for vision or multimodal workloads; strong computer vision and evaluation skills; BS/MS/PhD or equivalent experience.
featurization, embedding, inference, retrieval, vector search, approximate nearest neighbor retrieval, simulation workflows, synthetic data systems, computer vision models
2w
Save
Mark Applied
Hide
Senior Machine Learning Engineer, LLM Inference Optimization
Palo Alto or California
$195k-$262k/yr OnsiteFull Time
Nebius
NebiusNasdaq: NBIS: Builds cloud infrastructure and software for artificial intelligence development.
Expert Python and PyTorch skills, hands-on LLM/VLM inference deployment and optimization, knowledge of modern inference stacks, quantitative reasoning about latency/throughput/cost, and strong communication.
Python, PyTorch, vLLM, SGLang, TensorRT-LLM, Triton Inference Server, NVIDIA Dynamo, Ray Serve, KServe, CUDA, FlashInfer, LMCache, Ray
3w
Save
Mark Applied
Hide
Senior Software Engineer, AI Infrastructure - LVM Inference & Evaluation
Redwood City, California, United States
$168k-$205k/yr HybridFull Time
Ambient.ai
Ambient.ai: AI-powered physical security platform for proactive threat detection.
4+ YOE4+ years building infrastructure or production AI systems; strong Python; experience with ML infrastructure, LLM/LVM inference, inference optimization, evaluation frameworks, cloud and GPU workloads; BS/MS or equivalent.
Python, vLLM, Triton Inference Server, CUDA, NCCL, PyTorch, TensorRT, ONNX
2mo
Save
Mark Applied
Hide
ML Engineer - Inference & Model Deployment
Cupertino, California, United States
$250k-$310k/yr OnsiteFull Time
Hiring.Cafe
Hiring.Cafe: An AI-powered job search engine and aggregator.
Experience deploying and optimizing deep learning models in production, multi-GPU inference, profiling/benchmarking model performance, inference optimization techniques, and cloud/distributed systems familiarity.
vLLM, TensorRT, SGLang, GPU