12 gpu inference performance engineer jobs at 12 companies in Napa, CA
3d
Save
Mark Applied
Hide
3d
Inference Performance Engineer
San Francisco, California, United States
HybridFull Time
Adaption: Develops efficient AI systems that adapt in real-time.
5+ YOE5+ years in ML systems, inference infrastructure, or performance engineering; model-serving expertise; Python and systems-language proficiency; and GPU performance experience with measurable cost or latency improvements.
DigitalOceanNew York Stock Exchange: DOCN: Simplifies cloud infrastructure for developers, startups, and SMBs.
5+ YOE5+ years in high-performance computing or AI infrastructure, deep GPU and low-level optimization expertise, experience with CUDA/Triton/ROCm, distributed GPU parallelization, and system design for inference workloads.
TypeSafe AI: Building reliable, general frontier AI models for automation.
Deep CUDA/GPU kernel expertise, experience building and optimizing training and inference kernels, LLM training experience, profiling and eliminating performance bottlenecks.
Engram: Developing persistent memory layers for enterprise AI systems.
5+ YOE5+ years building training/inference systems; strong engineering skills; experience with ML frameworks, GPUs, distributed systems; bachelor's degree or equivalent experience.
Radical Numerics: Building general biological intelligence models for scientific discovery.
Deep expertise in large-model inference, GPU performance engineering, kernel development (CUDA/Triton), Python and PyTorch, distributed systems, and production model deployment.
Anthropic: Developing safe and reliable artificial intelligence systems.
Senior IC with deep systems or ML infrastructure experience, hands-on performance profiling and optimization, accelerator ecosystem expertise (CUDA/TPU/Trainium), strong software engineering and cross-org alignment skills, and a relevant bachelor’s degree or equivalent.
UnityNYSE: U: Provides software for creating real-time 3D interactive content.
5+ YOE5+ years in software/ML engineering with on-device or performance-critical systems; production deployment of transformer/diffusion models on-device; experience with inference runtimes, quantization, operator fusion, and GPU/compute APIs; strong Python; communication and mentoring skills.
Sail Research: Infrastructure platform for long-horizon agentic AI workloads.
Strong distributed systems fundamentals (concurrency, networking, databases, performance engineering). Ability to design, test, and reason about correctness in large-scale systems. Bonus: ML inference and GPU/accelerator experience.