27 gpu inference performance engineer jobs at 10 companies in Salinas, CA

2w
Save
Mark Applied
Hide
Inference Performance Engineer, AI Inference Configuration Optimization
Santa Clara or California
$124k-$242k/yr HybridFull Time
NVIDIA
NVIDIANASDAQ: NVDA: Designs graphics processing units and artificial intelligence hardware.
3+ YOEBS, MS, PhD, or equivalent experience; 3+ years engineering experience; AI inference optimization expertise; GPU profiling; Python; C++/CUDA; reproducible benchmarking and strong communication.
TensorRT-LLM, SGLang, vLLM, Dynamo, Nsight Systems, Nsight Compute, CUPTI, PyTorch profiler, Python, C++, CUDA, FlashInfer, NCCL, NIXL, NVSHMEM, MLPerf Inference, SemiAnalysis InferenceX
2w
Save
Mark Applied
Hide
Inference Performance Engineer, AI Inference Configuration Optimization
Santa Clara or United States
$124k-$242k/yr HybridFull Time
NVIDIA
NVIDIANASDAQ: NVDA: Designs GPU-accelerated computing and artificial intelligence hardware.
3+ YOEBachelor's, master's, or doctoral degree in a related field or equivalent experience; 3+ years' engineering experience; GPU profiling, Python, C++/CUDA, and AI inference optimization expertise required.
TensorRT-LLM, SGLang, vLLM, Dynamo, Nsight Systems, Nsight Compute, CUPTI, PyTorch profiler, Python, C++, CUDA, FlashInfer, NCCL, NIXL, NVSHMEM, MLPerf Inference, SemiAnalysis InferenceX
1w
Save
Mark Applied
Hide
Inference Systems Performance Architect
San Jose, California, United States
$245k-$325k/yr OnsiteFull Time
SambaNova Systems
SambaNova Systems: Develops custom AI hardware and software for enterprise computing.
12+ YOERequires 12+ years in performance engineering, distributed-systems analysis, workload generation, simulation, modeling, technical leadership, cross-functional influence, mentoring, and delivering complex ambiguous projects.
LLM, GPU, RDU, Headspace, Gympass+, One Medical, Employee Assistance Program (EAP)
2w
Save
Mark Applied
Hide
Backend Inference Runtime Engineer Graduate (AML Inference) - 2027 Start
San Jose, California, United States
OnsiteFull Time
ByteDance
ByteDance: Developing AI-driven content platforms and mobile applications.
Bachelor's or master's degree in a technical discipline; proficient in C/C++, Python, CUDA, GPU architecture, deep learning operators, inference compilation, performance analysis, and distributed model inference.
C, C++, Python, CUDA, Nsight, Profiler, vLLM, TensorRT-LLM, SGLang
1mo
Save
Mark Applied
Hide
Sr. Inference Optimization Engineer (local / edge runtime)
Santa Clara or Hillsboro or Folsom or Phoenix
$195k-$361k/yr HybridFull Time
Intel
IntelNasdaq: INTC: Designs and manufactures microprocessors and semiconductor components.
8+ YOE8+ years software development; strong C++ and/or Python; experience with LLM inference, profiling and optimizing CPU/GPU performance; Linux and low-level debugging expertise.
C++, Python, llama.cpp, vLLM, ggml, Vulkan, SYCL, oneAPI, CUDA, Metal, SIMD, Linux, GGUF, AWQ, GPTQ
2mo
Save
Mark Applied
Hide
ML Engineer - Inference & Model Deployment
Cupertino, California, United States
$250k-$310k/yr OnsiteFull Time
Hiring.Cafe
Hiring.Cafe: An AI-powered job search engine and aggregator.
Experience deploying and optimizing deep learning models in production, multi-GPU inference, profiling/benchmarking model performance, inference optimization techniques, and cloud/distributed systems familiarity.
vLLM, TensorRT, SGLang, GPU
1mo
Save
Mark Applied
Hide
Senior Software Engineer - LLM Inference
San Jose or Durham or Mexico City or Vancouver or Bengaluru or Pune or Hoofddorp or Belgrade or Barcelona or Singapore or Sydney or Tokyo
$171k-$257k/yr HybridFull Time
Nutanix
NutanixNASDAQ: NTNX: Sells cloud software and hyperconverged infrastructure for enterprises.
8+ YOE8+ years building distributed, high-performance systems; strong Go/Python, Docker, Kubernetes, CI/CD; knowledge of datacenter, OS internals, virtualization, and ML frameworks.
Docker, Kubernetes, Go, Python, CI/CD, TensorFlow, PyTorch, GPUs
1mo
Save
Mark Applied
Hide
AI Infra Engineer - Large Model Inference Systems (Multimodal/LLM/VLM)
San Jose, California, United States
$156k-$388k/yr OnsiteFull Time
TikTok
TikTok: Global short-form video hosting and social media platform.
2+ YOEBachelor's in CS or related,2+ years in high-performance computing or distributed scheduling,experience with large-model inference and system design,knowledge of asynchronous scheduling and resource pooling.
vLLM, SGLang, CUDA, Triton, Cutlass, PTQ, QAT
6d
Save
Mark Applied
Hide
Open Source Software Engineer — ML Systems & AMD Hardware
San Jose, California, United States
$179k-$306k/yr HybridFull Time
AMD
AMDNASDAQ: AMD: Designs and manufactures computer processors and graphics technology.
Systems software engineering experience with LLM inference, GPU kernel optimization, open source development, performance analysis, and GPU programming environments; bachelor's or master's degree preferred.
vLLM, SGLang, llama.cpp, ROCm, CUDA, Vulkan, SPIR-V
2mo
Save
Mark Applied
Hide
AI Infrastructure Engineer
San Jose, California, United States
$192k-$250k/yr OnsiteFull Time
NIO
NIONYSE: NIO: Designs and manufactures premium smart electric vehicles and technology
5+ YOE5+ years building and optimizing large-scale LLM/VLM inference systems; strong C/C++ and performance engineering skills; GPU/NPU programming (CUDA), PyTorch/TensorFlow, and BS/MS in CS/CE or related field required.
CUDA, PyTorch, TensorFlow, C/C++, AIOS
2w
Save
Mark Applied
Hide
Software Development Engineer, AI/ML, AWS Neuron, Model Inference
Cupertino, California, United States
$165k-$224k/yr OnsiteFull Time
Amazon
AmazonNASDAQ: AMZN: Global online retail and cloud computing technology provider.
3+ YOEBachelor's degree or equivalent; 3+ years professional software development and systems design experience; C++ or Python; machine learning, LLM, performance, memory, parallel computing, debugging, and profiling expertise.
AWS Neuron, Inferentia, Trainium, PyTorch, JAX, Python, C++, CUDA, CUTLASS, FlashInfer, Triton, vLLM, SGLang, TensorRT