40 gpu inference performance engineer jobs at 20 companies in Scotts Valley, CA

1mo
Save
Mark Applied
Hide
Senior GPU Inference Performance Engineer
Santa Clara, California, United States
$144k-$246k/yr HybridFull Time
AMD
AMDNASDAQ: AMD: Designs and manufactures computer processors and graphics technology.
Hands-on GPU performance engineering experience; proficent in GPU profiling, distributed inference, AI serving frameworks, Python and C/C++; bachelor's degree preferred.
ROCm, rocProfiler, Omniperf, Omnitrace, CUDA, Nsight Systems, Nsight Compute, DCGM, vLLM, SGLang, FP8, MXFP4, INT4, AWQ, GPTQ, RDMA, RoCE, NCCL, RCCL, Pensando, ConnectX, Docker, containerd, Kubernetes, Python, C++, C, HIP, perftest, ib_write_bw, tcpdump
2w
Save
Mark Applied
Hide
Inference Performance Engineer, AI Inference Configuration Optimization
Santa Clara or California
$124k-$242k/yr HybridFull Time
NVIDIA
NVIDIANASDAQ: NVDA: Designs graphics processing units and artificial intelligence hardware.
3+ YOEBS, MS, PhD, or equivalent experience; 3+ years engineering experience; AI inference optimization expertise; GPU profiling; Python; C++/CUDA; reproducible benchmarking and strong communication.
TensorRT-LLM, SGLang, vLLM, Dynamo, Nsight Systems, Nsight Compute, CUPTI, PyTorch profiler, Python, C++, CUDA, FlashInfer, NCCL, NIXL, NVSHMEM, MLPerf Inference, SemiAnalysis InferenceX
2w
Save
Mark Applied
Hide
Inference Performance Engineer, AI Inference Configuration Optimization
Santa Clara or United States
$124k-$242k/yr HybridFull Time
NVIDIA
NVIDIANASDAQ: NVDA: Designs GPU-accelerated computing and artificial intelligence hardware.
3+ YOEBachelor's, master's, or doctoral degree in a related field or equivalent experience; 3+ years' engineering experience; GPU profiling, Python, C++/CUDA, and AI inference optimization expertise required.
TensorRT-LLM, SGLang, vLLM, Dynamo, Nsight Systems, Nsight Compute, CUPTI, PyTorch profiler, Python, C++, CUDA, FlashInfer, NCCL, NIXL, NVSHMEM, MLPerf Inference, SemiAnalysis InferenceX
1w
Save
Mark Applied
Hide
Machine Learning Performance Engineer - Offboard Training & Inference
Sunnyvale or Washington, D.C. or San Diego or Fort Walton Beach or Ann Arbor or London or Stuttgart or Munich or Stockholm or Bangalore or Seoul or Tokyo
$215k-$285k/yr OnsiteFull Time
Applied Intuition
Applied Intuition: Developing software and simulation infrastructure for autonomous vehicles.
ML performance engineering experience with distributed training, batch inference, GPU or accelerator optimization, Python, and C++ or another systems language; strong debugging and analytical skills required.
FSDP, DeepSpeed, Megatron, NCCL, NVIDIA Triton Inference Server, TensorRT, ONNX Runtime, Ray, Python, C++, CUDA, Triton, CUTLASS, Nsight Systems, Nsight Compute, PyTorch Profiler, perf, Kubernetes, Slurm, ROS, OpenCV
2mo
Save
Mark Applied
Hide
Inference Optimization Engineer
Los Altos or United States or Canada
$198k-$286k/yr HybridFull Time
Modular
Modular: Unified software infrastructure and programming language for AI development.
5+ YOE5+ years in distributed systems or performance engineering; experience building reusable tooling; strong technical judgment, communication, and leadership; GPU/kernel, inference engine, Kubernetes, and LLM familiarity helpful.
GPU, ASIC, inference engine, Kubernetes, Modular Cloud, LLM architectures
1w
Save
Mark Applied
Hide
Inference Systems Performance Architect
San Jose, California, United States
$245k-$325k/yr OnsiteFull Time
SambaNova Systems
SambaNova Systems: Develops custom AI hardware and software for enterprise computing.
12+ YOERequires 12+ years in performance engineering, distributed-systems analysis, workload generation, simulation, modeling, technical leadership, cross-functional influence, mentoring, and delivering complex ambiguous projects.
LLM, GPU, RDU, Headspace, Gympass+, One Medical, Employee Assistance Program (EAP)
2w
Save
Mark Applied
Hide
Backend Inference Runtime Engineer Graduate (AML Inference) - 2027 Start
San Jose, California, United States
OnsiteFull Time
ByteDance
ByteDance: Developing AI-driven content platforms and mobile applications.
Bachelor's or master's degree in a technical discipline; proficient in C/C++, Python, CUDA, GPU architecture, deep learning operators, inference compilation, performance analysis, and distributed model inference.
C, C++, Python, CUDA, Nsight, Profiler, vLLM, TensorRT-LLM, SGLang
3mo
Save
Mark Applied
Hide
Performance Engineer
Palo Alto, California, United States
OnsiteFull Time
RadixArk
RadixArk: Building scalable open-source infrastructure for AI training and inference.
Strong systems engineering in performance-critical software; GPU/distributed systems; profiling tools; Python and C++; CUDA/Triton/ROCm/XLA familiarity; LLM inference concepts; ability to debug across software, hardware, and infra layers; strong communication.
CUDA, Triton, Pallas, ROCm, XLA, Python, C++
1mo
Save
Mark Applied
Hide
Sr. Inference Optimization Engineer (local / edge runtime)
Santa Clara or Hillsboro or Folsom or Phoenix
$195k-$361k/yr HybridFull Time
Intel
IntelNasdaq: INTC: Designs and manufactures microprocessors and semiconductor components.
8+ YOE8+ years software development; strong C++ and/or Python; experience with LLM inference, profiling and optimizing CPU/GPU performance; Linux and low-level debugging expertise.
C++, Python, llama.cpp, vLLM, ggml, Vulkan, SYCL, oneAPI, CUDA, Metal, SIMD, Linux, GGUF, AWQ, GPTQ
2mo
Save
Mark Applied
Hide
ML Engineer - Inference & Model Deployment
Cupertino, California, United States
$250k-$310k/yr OnsiteFull Time
Hiring.Cafe
Hiring.Cafe: An AI-powered job search engine and aggregator.
Experience deploying and optimizing deep learning models in production, multi-GPU inference, profiling/benchmarking model performance, inference optimization techniques, and cloud/distributed systems familiarity.
vLLM, TensorRT, SGLang, GPU
2w
Save
Mark Applied
Hide
AI Performance Modeling Engineer
Burlingame, California, United States
$180k-$225k/yr HybridFull Time
Quadric
Quadric: Designing licensable processor IP for on-device AI inference.
Strong Python and quantitative modeling skills; computer architecture knowledge; technical writing ability; AI inference or performance modeling expertise; BS, MS, PhD, or equivalent practical experience.
Python, C++, CUDA, Triton, gem5, Timeloop, MAESTRO, Accel-Sim
4d
Save
Mark Applied
Hide
Staff Software Engineer, Inference Performance Optimization, GenAI, DeepMind
Mountain View, California, United States
$207k-$300k/yr OnsiteFull Time
Google
GoogleNASDAQ: GOOGL: Provides online search, advertising, cloud computing, and consumer electronics.
8+ YOEBachelor's degree or equivalent experience, 8 years of software development, Python and C++, AI inference optimization expertise, and experience with serving codebases and performance tradeoffs.
Python, C++, vLLM, TensorRT-LLM, SGLang, Dynamo, PyTorch profiler, GPU, TPU
2mo
Save
Mark Applied
Hide
Software Engineer, ML Performance Optimization
Foster City, California, United States
$192k-$257k/yr OnsiteFull Time
Zoox
ZooxNASDAQ: AMZN: Developing autonomous robotaxis for urban ride-hailing services.
4+ YOE4+ years total exp; 2+ years in large-scale model training or inference; PyTorch; GPU-accelerated inference; profiling tools; Python or C++.
PyTorch, TensorRT, NVIDIA Nsight, Python, C++
1mo
Save
Mark Applied
Hide
Principal Machine Learning Engineer
Mountain View, California, United States
$278k-$417k/yr OnsiteFull Time
Unity
UnityNYSE: U: Provides software for creating real-time 3D interactive content.
8+ YOE4+ Mgmt8+ years software/ML engineering with 4+ years on-device/edge inference; production deployment of transformer/diffusion models; WebGPU/WGSL and GPU API performance tuning; proficiency with TypeScript/JavaScript and Python; leadership experience.
WebGPU, WebNN, WGSL, Metal, Vulkan, SPIR-V, D3D12, CUDA, Chrome, Dawn, PIX, Instruments, Snapdragon Profiler, Nsight, RenderDoc, ONNX Runtime Web, ONNX Runtime, Transformers.js, WebLLM, TensorFlow.js, CoreML, TFLite, ExecuTorch, TypeScript, JavaScript, Python, MLIR, TVM, IREE, XLA, wgpu
1mo
Save
Mark Applied
Hide
Senior Software Engineer - LLM Inference
San Jose or Durham or Mexico City or Vancouver or Bengaluru or Pune or Hoofddorp or Belgrade or Barcelona or Singapore or Sydney or Tokyo
$171k-$257k/yr HybridFull Time
Nutanix
NutanixNASDAQ: NTNX: Sells cloud software and hyperconverged infrastructure for enterprises.
8+ YOE8+ years building distributed, high-performance systems; strong Go/Python, Docker, Kubernetes, CI/CD; knowledge of datacenter, OS internals, virtualization, and ML frameworks.
Docker, Kubernetes, Go, Python, CI/CD, TensorFlow, PyTorch, GPUs
1mo
Save
Mark Applied
Hide
AI Infra Engineer - Large Model Inference Systems (Multimodal/LLM/VLM)
San Jose, California, United States
$156k-$388k/yr OnsiteFull Time
TikTok
TikTok: Global short-form video hosting and social media platform.
2+ YOEBachelor's in CS or related,2+ years in high-performance computing or distributed scheduling,experience with large-model inference and system design,knowledge of asynchronous scheduling and resource pooling.
vLLM, SGLang, CUDA, Triton, Cutlass, PTQ, QAT
1w
Save
Mark Applied
Hide
Member of Technical Staff, ML Inference Engineering
Palo Alto, California, United States
OnsiteFull Time
Sanas
Sanas: Provides real-time speech transformation and accent translation software.
5+ YOERequires 5+ years writing high-performance code, NVIDIA GPU and CUDA expertise, LLM serving knowledge, and research or systems experience in language or speech inference. Production-scale systems experience preferred.
CUDA, Triton, Kubernetes, InfiniBand, RoCE, LoRA
2mo
Save
Mark Applied
Hide
AI Infrastructure Engineer
San Jose, California, United States
$192k-$250k/yr OnsiteFull Time
NIO
NIONYSE: NIO: Designs and manufactures premium smart electric vehicles and technology
5+ YOE5+ years building and optimizing large-scale LLM/VLM inference systems; strong C/C++ and performance engineering skills; GPU/NPU programming (CUDA), PyTorch/TensorFlow, and BS/MS in CS/CE or related field required.
CUDA, PyTorch, TensorFlow, C/C++, AIOS
2w
Save
Mark Applied
Hide
Software Development Engineer, AI/ML, AWS Neuron, Model Inference
Cupertino, California, United States
$165k-$224k/yr OnsiteFull Time
Amazon
AmazonNASDAQ: AMZN: Global online retail and cloud computing technology provider.
3+ YOEBachelor's degree or equivalent; 3+ years professional software development and systems design experience; C++ or Python; machine learning, LLM, performance, memory, parallel computing, debugging, and profiling expertise.
AWS Neuron, Inferentia, Trainium, PyTorch, JAX, Python, C++, CUDA, CUTLASS, FlashInfer, Triton, vLLM, SGLang, TensorRT
1mo
Save
Mark Applied
Hide
[2026] Senior Machine Learning Engineer (Systems), Embodied AI/NPCs, ML Platform - PhD Early Career
San Mateo, California, United States
$197k-$243k/yr HybridFull Time
Roblox
RobloxNYSE: RBLX: Platform for creating and playing user-generated 3D digital experiences.
PhD (pursuing or completed) in a technical field; experience building end-to-end ML pipelines, model inference and deployment, distributed inference systems, Kubernetes and major cloud providers (AWS/Azure/GCP); strong systems and performance optimization skills.
Kubernetes, AWS, Azure, GCP, GPU, LLMs, Roblox Studio IDE