20 gpu performance profiling engineer jobs at 16 companies in Vallejo, CA
1mo
Save
Mark Applied
Hide
1mo
GPU Kernel Engineer
San Francisco, California, United States
$180k-$280k/yrOnsiteFull Time
TypeSafe AI: Private frontier AI lab building reliable, general AI systems for real-world automation and decision-making.
Deep CUDA/GPU kernel expertise, experience building and optimizing training and inference kernels, LLM training experience, profiling and eliminating performance bottlenecks.
RadixArk: Private AI infrastructure building open training, inference, and post-training systems for developers and research labs.
Strong systems engineering in performance-critical software; GPU/distributed systems; profiling tools; Python and C++; CUDA/Triton/ROCm/XLA familiarity; LLM inference concepts; ability to debug across software, hardware, and infra layers; strong communication.
Member of Technical Staff, Hardware, Performance Engineer
Palo Alto or Austin
$200k-$420k/yrOnsiteFull Time
River AI: Full-stack AI building personal AI, custom models, training infrastructure, and local hardware for developers and enterprises.
5+ YOEBachelor's in EE/CE/CS and 5+ years experience with advanced process nodes; expert in C/C++ or SystemC; familiarity with compilers (LLVM/GCC/XLA), CUDA/Triton, computer architecture, and profiling/trace analysis.
Research Scientist / Engineer – Performance Optimization
Redwood City, California, United States
OnsiteFull Time
Luma AI: AI is a private creative AI platform generating video and images for creators and teams.
Expert GPU/CPU/accelerator optimization with Triton/CUDA, strong PyTorch and kernel development, profiling tools experience, deep transformer knowledge, and distributed deployment skills.
Staff Software Engineer - Video Performance - (Bay area only)
San Francisco, California, United States
$251k-$329k/yrOnsiteFull Time
Canva: Online graphic design and visual communication platform.
Strong C++ proficiency; experience optimizing multithreaded systems, graphics APIs, multimedia, profiling, telemetry, and CPU/GPU architectures. Systems language experience and performance engineering expertise are valued.
Ollama: Open-model software platform helping developers run models locally and in the cloud.
Experience with systems programming (Go,C,C++), GPU or low-level performance work, profiling and optimizing real workloads, and shipping software across macOS, Linux, and Windows.
Genesis AI: Private French full-stack robotics building general-purpose robots and physical-AI systems for factories, laboratories, hospitals, and homes.
Strong Python engineering, large-codebase architecture, developer tooling or CI, performance profiling, physics and robotics knowledge, and familiarity with JIT compilation or GPU computing.
Google DeepMind: Google DeepMind is a private AI research laboratory developing safe artificial intelligence systems for Alphabet.
2+ YOEBachelor’s degree or equivalent practical experience; 2+ years developing in C++ or Python; profiling tools; hardware accelerator software development; memory-hierarchy or instruction-level tuning.
Zensors: AI technology helping airports, retailers, and other large physical businesses automate operations with spatial intelligence.
BS/MS/PhD in CS or EE; strong C/C++ and Python; model optimization, quantization, pruning; GPU performance tuning; profiling tools; cross-functional collaboration.
MetaNASDAQ: META: Builds technologies that help people connect, find communities, and grow businesses.
6+ YOEBachelor's in CS/CE or equivalent, 6+ years software engineering experience in ML systems or HPC, proficiency with PyTorch/TensorFlow, C++ and Python, distributed systems, and performance profiling.
Anthropic: AI research developing safe and steerable AI systems.
Senior IC with deep systems or ML infrastructure experience, hands-on performance profiling and optimization, accelerator ecosystem expertise (CUDA/TPU/Trainium), strong software engineering and cross-org alignment skills, and a relevant bachelor’s degree or equivalent.
Unconventional AI: Private AI hardware startup building energy-efficient computing substrates for artificial-intelligence workloads.
MS/PhD (or equivalent) in quantitative field, deep practical experience with ML stack and GPU performance optimization, proficiency in profiling and optimizing large ML codebases.
Principal Physics Programmer, C++ (Modeling & Simulation)
San Francisco, California, United States
HybridFull Time
Code Metal: AI software translating, verifying, and optimizing code for mission-critical industries deploying software on specialized hardware.
8+ YOE8+ years professional C++ development building performance-critical real-time systems; experience with game-engine or engine-level systems, data-oriented design, concurrency, profiling, and optimization; eligibility for Secret clearance.