14 gpu inference performance engineer jobs at 8 companies in Pacific Grove, CA

1w
Save
Mark Applied
Hide
Inference Systems Performance Architect
San Jose, California, United States
$245k-$325k/yr OnsiteFull Time
SambaNova Systems
SambaNova Systems: Develops custom AI hardware and software for enterprise computing.
12+ YOERequires 12+ years in performance engineering, distributed-systems analysis, workload generation, simulation, modeling, technical leadership, cross-functional influence, mentoring, and delivering complex ambiguous projects.
LLM, GPU, RDU, Headspace, Gympass+, One Medical, Employee Assistance Program (EAP)
2w
Save
Mark Applied
Hide
Backend Inference Runtime Engineer Graduate (AML Inference) - 2027 Start
San Jose, California, United States
OnsiteFull Time
ByteDance
ByteDance: Developing AI-driven content platforms and mobile applications.
Bachelor's or master's degree in a technical discipline; proficient in C/C++, Python, CUDA, GPU architecture, deep learning operators, inference compilation, performance analysis, and distributed model inference.
C, C++, Python, CUDA, Nsight, Profiler, vLLM, TensorRT-LLM, SGLang
2mo
Save
Mark Applied
Hide
ML Engineer - Inference & Model Deployment
Cupertino, California, United States
$250k-$310k/yr OnsiteFull Time
Hiring.Cafe
Hiring.Cafe: An AI-powered job search engine and aggregator.
Experience deploying and optimizing deep learning models in production, multi-GPU inference, profiling/benchmarking model performance, inference optimization techniques, and cloud/distributed systems familiarity.
vLLM, TensorRT, SGLang, GPU
2w
Save
Mark Applied
Hide
Backend Inference Framework Engineer Graduate (AML Inference) - 2027 Start
San Jose, California, United States
OnsiteFull Time
ByteDance
ByteDance: Developing AI-driven content platforms and mobile applications.
Bachelor's or master's degree in a technical discipline; C/C++, Linux, data structures, multithreaded concurrency, distributed services, performance optimization, and strong communication skills required.
Linux, C, C++, Redis, RocksDB, BRPC, GRPC, GPU
1mo
Save
Mark Applied
Hide
Senior Software Engineer - LLM Inference
San Jose or Durham or Mexico City or Vancouver or Bengaluru or Pune or Hoofddorp or Belgrade or Barcelona or Singapore or Sydney or Tokyo
$171k-$257k/yr HybridFull Time
Nutanix
NutanixNASDAQ: NTNX: Sells cloud software and hyperconverged infrastructure for enterprises.
8+ YOE8+ years building distributed, high-performance systems; strong Go/Python, Docker, Kubernetes, CI/CD; knowledge of datacenter, OS internals, virtualization, and ML frameworks.
Docker, Kubernetes, Go, Python, CI/CD, TensorFlow, PyTorch, GPUs
2mo
Save
Mark Applied
Hide
Software Engineer 2 - LLM Inference
Vancouver or San Jose or Durham or Mexico City or Bengaluru or Pune or Hoofddorp or Belgrade or Barcelona or Singapore or Sydney or Tokyo
$129k-$193k/yr HybridFull Time
Nutanix
NutanixNASDAQ: NTNX: Sells cloud software and hyperconverged infrastructure for enterprises.
2+ YOE2–5 years product development experience; strong programming fundamentals; experience with Docker, Kubernetes, Go, Python, CI/CD, distributed systems, datacenter design, OS internals, and high-performance system tuning; Bachelor's/Masters in CS or equivalent.
Docker, Kubernetes, Go, Python, CI/CD, TensorFlow, PyTorch, Nutanix Cloud Platform for AI, GPUs
1mo
Save
Mark Applied
Hide
AI Infra Engineer - Large Model Inference Systems (Multimodal/LLM/VLM)
San Jose, California, United States
$156k-$388k/yr OnsiteFull Time
TikTok
TikTok: Global short-form video hosting and social media platform.
2+ YOEBachelor's in CS or related,2+ years in high-performance computing or distributed scheduling,experience with large-model inference and system design,knowledge of asynchronous scheduling and resource pooling.
vLLM, SGLang, CUDA, Triton, Cutlass, PTQ, QAT
1mo
Save
Mark Applied
Hide
Senior AI Infra Engineer - Large Model Inference Systems (Multimodal/LLM/VLM)
San Jose, California, United States
$213k-$450k/yr OnsiteFull Time
TikTok
TikTok: Global short-form video hosting and social media platform.
4+ YOEBachelor's degree,4+ years in high-performance computing or distributed scheduling, familiarity with large-model architectures, strong system design and performance-optimization skills, experience with CUDA/Triton/Cutlass and inference frameworks.
CUDA, Triton, Cutlass, vLLM, SGLang
6d
Save
Mark Applied
Hide
Open Source Software Engineer — ML Systems & AMD Hardware
San Jose, California, United States
$179k-$306k/yr HybridFull Time
AMD
AMDNASDAQ: AMD: Designs and manufactures computer processors and graphics technology.
Systems software engineering experience with LLM inference, GPU kernel optimization, open source development, performance analysis, and GPU programming environments; bachelor's or master's degree preferred.
vLLM, SGLang, llama.cpp, ROCm, CUDA, Vulkan, SPIR-V
2mo
Save
Mark Applied
Hide
AI Infrastructure Engineer
San Jose, California, United States
$192k-$250k/yr OnsiteFull Time
NIO
NIONYSE: NIO: Designs and manufactures premium smart electric vehicles and technology
5+ YOE5+ years building and optimizing large-scale LLM/VLM inference systems; strong C/C++ and performance engineering skills; GPU/NPU programming (CUDA), PyTorch/TensorFlow, and BS/MS in CS/CE or related field required.
CUDA, PyTorch, TensorFlow, C/C++, AIOS
2w
Save
Mark Applied
Hide
Software Development Engineer, AI/ML, AWS Neuron, Model Inference
Cupertino, California, United States
$165k-$224k/yr OnsiteFull Time
Amazon
AmazonNASDAQ: AMZN: Global online retail and cloud computing technology provider.
3+ YOEBachelor's degree or equivalent; 3+ years professional software development and systems design experience; C++ or Python; machine learning, LLM, performance, memory, parallel computing, debugging, and profiling expertise.
AWS Neuron, Inferentia, Trainium, PyTorch, JAX, Python, C++, CUDA, CUTLASS, FlashInfer, Triton, vLLM, SGLang, TensorRT
1w
Save
Mark Applied
Hide
Senior Software Development Engineer, AI/ML, AWS Neuron, Model Inference
Cupertino, California, United States
$193k-$262k/yr OnsiteFull Time
Amazon
AmazonNASDAQ: AMZN: Global online retail and cloud computing technology provider.
5+ YOEBachelor's degree and 5+ years of professional software development and systems design experience. Requires C++ or Python, machine learning and LLM knowledge, system performance, memory management, parallel computing, debugging, and profiling.
Amazon Neuron, PyTorch, JAX, Python, C++, CUDA, CUTLASS, FlashInfer, Triton, vLLM, SGLang, TensorRT
1mo
Save
Mark Applied
Hide
Production Engineer - Applied Machine Learning
San Jose, California, United States
OnsiteFull Time
ByteDance
ByteDance: Developing AI-driven content platforms and mobile applications.
Bachelor's degree in CS or related field, proficiency with Linux and scripting (Shell, Python, Go, or C++), experience with ML training/inference architectures, Kubernetes and GPU clusters, online troubleshooting, performance analysis, and automation.
Linux, Shell, Python, Go, C++, Kubernetes, NoSQL, CI/CD, FinOps
3w
Save
Mark Applied
Hide
AI Infrastructure Engineer Graduate (Algorithm Infrastructure) - 2027 Start (PhD)
San Jose, California, United States
$156k-$388k/yr OnsiteFull Time
TikTok
TikTok: Global short-form video hosting and social media platform.
PhD in CS or related field, strong system design for high-concurrency distributed inference, familiarity with large-model architectures, performance optimization skills, and experience with CUDA and Triton.
CUDA, Triton