16 gpu inference performance engineer jobs at 16 companies in Vallejo, CA
3d
Save
Mark Applied
Hide
3d
Inference Performance Engineer
San Francisco, California, United States
HybridFull Time
Adaption: Develops efficient AI systems that adapt in real-time.
5+ YOE5+ years in ML systems, inference infrastructure, or performance engineering; model-serving expertise; Python and systems-language proficiency; and GPU performance experience with measurable cost or latency improvements.
DigitalOceanNew York Stock Exchange: DOCN: Simplifies cloud infrastructure for developers, startups, and SMBs.
5+ YOE5+ years in high-performance computing or AI infrastructure, deep GPU and low-level optimization expertise, experience with CUDA/Triton/ROCm, distributed GPU parallelization, and system design for inference workloads.
TypeSafe AI: Building reliable, general frontier AI models for automation.
Deep CUDA/GPU kernel expertise, experience building and optimizing training and inference kernels, LLM training experience, profiling and eliminating performance bottlenecks.
RadixArk: Building scalable open-source infrastructure for AI training and inference.
Strong systems engineering in performance-critical software; GPU/distributed systems; profiling tools; Python and C++; CUDA/Triton/ROCm/XLA familiarity; LLM inference concepts; ability to debug across software, hardware, and infra layers; strong communication.
Engram: Developing persistent memory layers for enterprise AI systems.
5+ YOE5+ years building training/inference systems; strong engineering skills; experience with ML frameworks, GPUs, distributed systems; bachelor's degree or equivalent experience.
Radical Numerics: Building general biological intelligence models for scientific discovery.
Deep expertise in large-model inference, GPU performance engineering, kernel development (CUDA/Triton), Python and PyTorch, distributed systems, and production model deployment.
Anthropic: Developing safe and reliable artificial intelligence systems.
Senior IC with deep systems or ML infrastructure experience, hands-on performance profiling and optimization, accelerator ecosystem expertise (CUDA/TPU/Trainium), strong software engineering and cross-org alignment skills, and a relevant bachelor’s degree or equivalent.
Member of Technical Staff, ML Inference Engineering
Palo Alto, California, United States
OnsiteFull Time
Sanas: Provides real-time speech transformation and accent translation software.
5+ YOERequires 5+ years writing high-performance code, NVIDIA GPU and CUDA expertise, LLM serving knowledge, and research or systems experience in language or speech inference. Production-scale systems experience preferred.
UnityNYSE: U: Provides software for creating real-time 3D interactive content.
5+ YOE5+ years in software/ML engineering with on-device or performance-critical systems; production deployment of transformer/diffusion models on-device; experience with inference runtimes, quantization, operator fusion, and GPU/compute APIs; strong Python; communication and mentoring skills.
[2026] Senior Machine Learning Engineer (Systems), Embodied AI/NPCs, ML Platform - PhD Early Career
San Mateo, California, United States
$197k-$243k/yrHybridFull Time
RobloxNYSE: RBLX: Platform for creating and playing user-generated 3D digital experiences.
PhD (pursuing or completed) in a technical field; experience building end-to-end ML pipelines, model inference and deployment, distributed inference systems, Kubernetes and major cloud providers (AWS/Azure/GCP); strong systems and performance optimization skills.
Kubernetes, AWS, Azure, GCP, GPU, LLMs, Roblox Studio IDE
Sail Research: Infrastructure platform for long-horizon agentic AI workloads.
Strong distributed systems fundamentals (concurrency, networking, databases, performance engineering). Ability to design, test, and reason about correctness in large-scale systems. Bonus: ML inference and GPU/accelerator experience.