55 inference optimization engineer jobs at 16 companies in Salinas, CA

2w
Save
Mark Applied
Hide
Inference Performance Engineer, AI Inference Configuration Optimization
Santa Clara or California
$124k-$242k/yr HybridFull Time
NVIDIA
NVIDIANASDAQ: NVDA: Designs graphics processing units and artificial intelligence hardware.
3+ YOEBS, MS, PhD, or equivalent experience; 3+ years engineering experience; AI inference optimization expertise; GPU profiling; Python; C++/CUDA; reproducible benchmarking and strong communication.
TensorRT-LLM, SGLang, vLLM, Dynamo, Nsight Systems, Nsight Compute, CUPTI, PyTorch profiler, Python, C++, CUDA, FlashInfer, NCCL, NIXL, NVSHMEM, MLPerf Inference, SemiAnalysis InferenceX
2w
Save
Mark Applied
Hide
Inference Performance Engineer, AI Inference Configuration Optimization
Santa Clara or United States
$124k-$242k/yr HybridFull Time
NVIDIA
NVIDIANASDAQ: NVDA: Designs GPU-accelerated computing and artificial intelligence hardware.
3+ YOEBachelor's, master's, or doctoral degree in a related field or equivalent experience; 3+ years' engineering experience; GPU profiling, Python, C++/CUDA, and AI inference optimization expertise required.
TensorRT-LLM, SGLang, vLLM, Dynamo, Nsight Systems, Nsight Compute, CUPTI, PyTorch profiler, Python, C++, CUDA, FlashInfer, NCCL, NIXL, NVSHMEM, MLPerf Inference, SemiAnalysis InferenceX
1mo
Save
Mark Applied
Hide
Sr. Inference Optimization Engineer (local / edge runtime)
Santa Clara or Hillsboro or Folsom or Phoenix
$195k-$361k/yr HybridFull Time
Intel
IntelNasdaq: INTC: Designs and manufactures microprocessors and semiconductor components.
8+ YOE8+ years software development; strong C++ and/or Python; experience with LLM inference, profiling and optimizing CPU/GPU performance; Linux and low-level debugging expertise.
C++, Python, llama.cpp, vLLM, ggml, Vulkan, SYCL, oneAPI, CUDA, Metal, SIMD, Linux, GGUF, AWQ, GPTQ
1mo
Save
Mark Applied
Hide
Principal LLM Inference Engineer
Santa Clara, California, United States
$195k-$285k/yr HybridFull Time
d-Matrix: Develops high-performance semiconductor chips for generative AI inference.
10+ YOEBachelor's in CS/EE (or equivalent) with 10+ years experience (Master/PhD with 6+ years preferred); strong Python and C/C++; experience optimizing LLM inference, quantization, batching, GPU kernel programming and contributor-level work on inference frameworks.
Python, C, C++, vLLM, SGLang, TensorRT-LLM, ONNX Runtime, CUDA, Triton, JAX
1mo
Save
Mark Applied
Hide
AI Inference Engineer - Speech
Seattle or San Jose
$152k-$332k/yr HybridFull Time
Zoom
ZoomNasdaq: ZM: Provides a cloud-based platform for video, voice, and collaboration.
3+ YOEMaster's in CS/EE or related,3+ years in speech recognition or model inference,deep learning expertise,experience with Python,C/C++,CUDA,TensorRT,PyTorch,TensorFlow and GPU optimization.
Python, shell, C/C++, PyTorch, TensorFlow, CUDA, TensorRT, CUDA Graphs, NVIDIA GPUs, TPU, BrightHire
1w
Save
Mark Applied
Hide
Senior/Sr. Staff AI Infrastructure Engineer, Inference & Optimization
San Jose or Mountain View
$170k-$351k/yr OnsiteFull Time
DiDi Autonomous Driving
DiDi Autonomous Driving: Develops Level 4 autonomous driving technology for shared mobility.
3+ YOEMaster's degree or higher in a technical field; 3+ years in HPC, AI infrastructure, model optimization, or embedded deployment; C++, Python, CUDA, OpenMP, inference engines, GPU architectures, and system profiling expertise.
C++, Python, CUDA, OpenMP, TensorRT, ONNX Runtime, vLLM, SGLang, TensorRT-LLM, NVIDIA Hopper, NVIDIA Thor, TGI, LightLLM, PyTorch, INT8, FP8, AWQ, LLaMA, Qwen, GPT
1mo
Save
Mark Applied
Hide
Research Engineer - LLM/VLM Inference Optimization (Seed Infra)
San Jose, California, United States
OnsiteFull Time
ByteDance
ByteDance: Developing AI-driven content platforms and mobile applications.
Bachelor's in CS/EE/Software or related; strong C/C++ and Python; experience with PyTorch or TensorFlow; production LLM/VLM inference optimization experience; familiarity with GPU architecture and containerized server debugging.
C/C++, Python, PyTorch, TensorFlow, CUDA, OpenCL, TensorRT, Triton, CUTLASS, FlashAttention
2mo
Save
Mark Applied
Hide
ML Engineer - Inference & Model Deployment
Cupertino, California, United States
$250k-$310k/yr OnsiteFull Time
Hiring.Cafe
Hiring.Cafe: An AI-powered job search engine and aggregator.
Experience deploying and optimizing deep learning models in production, multi-GPU inference, profiling/benchmarking model performance, inference optimization techniques, and cloud/distributed systems familiarity.
vLLM, TensorRT, SGLang, GPU
1w
Save
Mark Applied
Hide
Sr. Machine Learning Engineer, Foundation Models Inference - Cloud OS & Inference
Santa Clara, California, United States
OnsiteFull Time
Apple
AppleNASDAQ: AAPL: Designs and sells consumer electronics, software, and online services.
Build and optimize inference frameworks, services, and tools for large-scale foundation models across cloud infrastructure, including language, vision, and speech models.
2mo
Save
Mark Applied
Hide
Sr. Lead AI Engineer (Inference Optimization, FM hosting, AI Platform)
San Jose or San Francisco or New York City or Cambridge or McLean
$230k-$286k/yr OnsiteFull Time
Capital One
Capital OneNYSE: COF: Provides credit card, banking, and auto loan services.
6+ YOEBachelor's plus 6 years or master's plus 4 years developing AI/ML technologies, and 6 years programming with Python, Go, Scala, or Java. Cloud AI deployment and team leadership are preferred.
AWS Ultraclusters, Hugging Face, VectorDBs, NeMo Guardrails, PyTorch, Python, Go, Scala, Java, AWS, Google Cloud, Azure, C++, C#, Golang
1mo
Save
Mark Applied
Hide
Senior Lead AI Engineer (FM Hosting, LLM Inference)
New York or McLean or Cambridge or San Jose
$251k-$286k/yr OnsiteFull Time
Capital One
Capital OneNYSE: COF: Financial services offering credit cards, banking, and loans.
6+ YOEBachelor's in CS/AI/EE/CE or related with 6+ years (or Master's with 4+ years); 6+ years programming with Python, Go, Scala, or Java; experience deploying scalable AI systems, LLM inference, similarity search, and optimization of training/inference.
AWS Ultraclusters, Huggingface, VectorDBs, Nemo Guardrails, PyTorch, Python, Go, Scala, Java, C++, C#, Golang, AWS, Google Cloud, Azure
1w
Save
Mark Applied
Hide
Sr. Software Development Engineer, Inference Team - AWS Neuron
Seattle or Cupertino or Los Angeles County
$168k-$262k/yr OnsiteFull Time
Amazon
AmazonNASDAQ: AMZN: Global online retail and cloud computing technology provider.
5+ YOERequires 5+ years of software development and programming experience, a computer science degree or equivalent, machine learning optimization experience, and knowledge of inference frameworks and accelerator hardware.
vLLM, SGLang, Java, C++, C#, PyTorch, JAX, AWS Neuron, CUDA, Triton
2mo
Save
Mark Applied
Hide
Sr. Lead AI Engineer (Inference Optimization, FM hosting, AI Platform)
San Jose or San Francisco or New York City or Cambridge or McLean
$230k-$286k/yr OnsiteFull Time
Capital One
Capital OneNYSE: COF: A diversified financial services providing banking and credit products.
6+ YOEBachelor's degree plus 6 years or master's degree plus 4 years developing AI/ML technologies; 6 years programming with Python, Go, Scala, or Java; cloud AI deployment experience preferred.
AWS Ultraclusters, Hugging Face, VectorDBs, Nemo Guardrails, PyTorch, Python, Go, Scala, Java, Google Cloud, Azure, C++, C#, Golang
1mo
Save
Mark Applied
Hide
Senior AI Infra Engineer - Large Model Inference Systems (Multimodal/LLM/VLM)
San Jose, California, United States
$213k-$450k/yr OnsiteFull Time
TikTok
TikTok: Global short-form video hosting and social media platform.
4+ YOEBachelor's degree,4+ years in high-performance computing or distributed scheduling, familiarity with large-model architectures, strong system design and performance-optimization skills, experience with CUDA/Triton/Cutlass and inference frameworks.
CUDA, Triton, Cutlass, vLLM, SGLang
2mo
Save
Mark Applied
Hide
Distinguished Engineer - AI (San Jose, CA, US, 95128)
San Jose, California, United States
$266k-$396k/yr OnsiteFull Time
NetApp
NetAppNASDAQ: NTAP: Sells enterprise data storage and cloud management software.
15+ YOE15+ years building low-latency, fault-tolerant distributed systems and AI/ML inference platforms; expertise with inference engines, model optimization, storage for AI, RDMA/DPDK, and Kubernetes-based orchestration.
TensorRT, vLLM, ONNX Runtime, Triton, RDMA, DPDK, Kubernetes
2mo
Save
Mark Applied
Hide
Distinguished Engineer - AI
San Jose, California, United States
$266k-$396k/yr OnsiteFull Time
NetApp
NetAppNasdaq: NTAP: Provides intelligent data infrastructure for hybrid cloud environments.
15+ YOEExpert in AI inferencing and distributed systems at scale with 15+ years experience; hands-on with inference engines, model optimization, GPU/TPU orchestration, Kubernetes, RDMA/DPDK; strong architecture, communication, and mentorship skills.
TensorRT, vLLM, ONNX Runtime, Triton, RDMA, DPDK, Kubernetes, GPU, TPU
2mo
Save
Mark Applied
Hide
ML Systems Engineer
San Jose or Santa Barbara
$150k-$350k/yr OnsiteFull Time
Alpha Design AI
Alpha Design AI: AI-native EDA platform for semiconductor design and verification.
Experience with large-scale ML systems and GPU computing; strong Python and C++/CUDA skills; familiarity with vLLM, PyTorch, SGLang, Ray; experience deploying and optimizing LLMs, profiling and benchmarking inference.
Python, C++, CUDA, SGLang, vLLM, PyTorch, Ray
1mo
Save
Mark Applied
Hide
Super Sparks-校招-高性能AI推理平台研究员/工程师
Shanghai or San Jose
OnsiteFull Time
Nio
NioNYSE: NIO: Designs and manufactures smart premium electric vehicles.
Master's degree or higher in CS/AI/EE or related; deep learning framework proficiency (PyTorch, TensorFlow, ONNX); model optimization and edge deployment experience; embedded Linux/RTOS/QNX and C/C++ skills; experience with heterogeneous platforms (NPU/DSP/GPU) preferred.
PyTorch, TensorFlow, ONNX, TensorRT, llama.cpp, MNN, TNN, C/C++, QNX, ARM
2mo
Save
Mark Applied
Hide
AI Infrastructure Engineer
San Jose, California, United States
$192k-$250k/yr OnsiteFull Time
NIO
NIONYSE: NIO: Designs and manufactures premium smart electric vehicles and technology
5+ YOE5+ years building and optimizing large-scale LLM/VLM inference systems; strong C/C++ and performance engineering skills; GPU/NPU programming (CUDA), PyTorch/TensorFlow, and BS/MS in CS/CE or related field required.
CUDA, PyTorch, TensorFlow, C/C++, AIOS
1mo
Save
Mark Applied
Hide
Principal Scientist - Data Pipeline Engineer
San Jose or Seattle or San Francisco
$206k-$388k/yr OnsiteFull Time
Adobe
AdobeNASDAQ: ADBE: Provides software for digital media creation and marketing analytics
10+ YOE10+ years in data engineering/ML infrastructure, distributed systems expertise, Python and a systems language, experience with Ray or Spark, GPU inference optimization, large-scale databases and data curation for model training.
Ray, Spark, Python, C++, Rust, Go, Java