10 gpu performance profiling engineer jobs at 8 companies in Napa, CA

1mo
Save
Mark Applied
Hide
GPU Kernel Engineer
San Francisco, California, United States
$180k-$280k/yr OnsiteFull Time
TypeSafe AI
TypeSafe AI: Private frontier AI lab building reliable, general AI systems for real-world automation and decision-making.
Deep CUDA/GPU kernel expertise, experience building and optimizing training and inference kernels, LLM training experience, profiling and eliminating performance bottlenecks.
CUDA, CuTe DSL
2mo
Save
Mark Applied
Hide
Member of Technical Staff, TPU & AMD GPU Performance Engineering
San Francisco, California, United States
$200k-$400k/yr OnsiteFull Time
Inferact
Inferact: Private AI infrastructure building open-source software that makes LLM inference faster and cheaper.
Bachelor's or equivalent experience in CS/engineering; hands-on AMD GPU/TPU optimization; experience with ROCm/HIP/Triton, XLA/JAX/Pallas; strong profiling, benchmarking, and cross-platform performance skills.
vLLM, ROCm, HIP, Triton, CK, AITER, TPU, XLA, JAX, Pallas, SGLang, TensorRT-LLM, ATOM, MLIR, LLVM, PyTorch
2mo
Save
Mark Applied
Hide
Member of Technical Staff, AMD GPU Performance Engineering
San Francisco, California, United States
$200k-$400k/yr OnsiteFull Time
Inferact
Inferact: Private AI infrastructure building open-source software that makes LLM inference faster and cheaper.
Bachelor's or equivalent experience; hands-on AMD GPU optimization using ROCm/HIP/Triton/CK/AITER; deep understanding of AMD GPU execution, memory, toolchains; experience optimizing ML kernels and strong profiling/benchmarking skills.
ROCm, HIP, Triton, CK, AITER, vLLM, SGLang, TensorRT-LLM, MLIR, LLVM, PyTorch
1w
Save
Mark Applied
Hide
Staff Software Engineer - Video Performance - (Bay area only)
San Francisco, California, United States
$251k-$329k/yr OnsiteFull Time
Canva
Canva: Online graphic design and visual communication platform.
Strong C++ proficiency; experience optimizing multithreaded systems, graphics APIs, multimedia, profiling, telemetry, and CPU/GPU architectures. Systems language experience and performance engineering expertise are valued.
C++, Rust, GLSL, HLSL, Metal, Vulkan, WebGPU, OpenGL, Perf, Instruments, Chrome DevTools, Systrace, H.264, H.265, VP9, AV1, iOS, Android, Web, SIMD
1w
Save
Mark Applied
Hide
Staff Software Engineer - Video Performance - (Bay area only)
San Francisco, California, United States
$251k-$329k/yr OnsiteFull Time
Canva
Canva: Online graphic design and visual communication platform.
Strong C++ proficiency; systems performance optimization, CPU/GPU architecture, SIMD, graphics APIs, multimedia codecs, profiling, diagnostics, telemetry, and cross-team technical collaboration experience.
C++, Rust, GLSL, HLSL, Metal, Vulkan, WebGPU, OpenGL, Perf, Instruments, Chrome DevTools, Systrace, CMake, Web, iOS, Android, H.264, H.265, VP9, AV1, SIMD
1w
Save
Mark Applied
Hide
Software Engineer, Hardware Accelerators Performance, GeminiApp, DeepMind
Mountain View or San Francisco
$147k-$210k/yr OnsiteFull Time
Google DeepMind
Google DeepMind: Google DeepMind is a private AI research laboratory developing safe artificial intelligence systems for Alphabet.
2+ YOEBachelor’s degree or equivalent practical experience; 2+ years developing in C++ or Python; profiling tools; hardware accelerator software development; memory-hierarchy or instruction-level tuning.
C++, Python, gProf, Valgrind, VTune, CPUs, GPUs, TPUs, JAX
3mo
Save
Mark Applied
Hide
AI/ML Infrastructure Engineer
San Francisco, California, United States
OnsiteFull Time
Zensors
Zensors: AI technology helping airports, retailers, and other large physical businesses automate operations with spatial intelligence.
BS/MS/PhD in CS or EE; strong C/C++ and Python; model optimization, quantization, pruning; GPU performance tuning; profiling tools; cross-functional collaboration.
C/C++, Python, Nsight Systems, Nsight Compute, PyTorch, CUDA, TensorRT, NVIDIA DeepStream, DALI, FFmpeg, TVM, MLIR, ONNX Runtime, Triton
2mo
Save
Mark Applied
Hide
Staff+ Software Engineer, Inference Runtime
San Francisco or Seattle or New York City
$405k-$485k/yr HybridFull Time
Anthropic
Anthropic: AI research developing safe and steerable AI systems.
Senior IC with deep systems or ML infrastructure experience, hands-on performance profiling and optimization, accelerator ecosystem expertise (CUDA/TPU/Trainium), strong software engineering and cross-org alignment skills, and a relevant bachelor’s degree or equivalent.
Rust, Python, CUDA, XLA, Triton, NeuronX, AWS Neuron, Kubernetes, CI/CD
3mo
Save
Mark Applied
Hide
Systems Software Engineer (C++)
San Francisco, California, United States
OnsiteFull Time
Thunder Compute
Thunder Compute: GPU virtualization serving developers, enterprises, and AI innovators with cloud infrastructure.
Expert C++ performance optimization, low-level networking, and production systems ownership.
C++, Low-level networking, GPU optimization, Compilers, Performance profiling
2mo
Save
Mark Applied
Hide
Principal Physics Programmer, C++ (Modeling & Simulation)
San Francisco, California, United States
HybridFull Time
Code Metal
Code Metal: AI software translating, verifying, and optimizing code for mission-critical industries deploying software on specialized hardware.
8+ YOE8+ years professional C++ development building performance-critical real-time systems; experience with game-engine or engine-level systems, data-oriented design, concurrency, profiling, and optimization; eligibility for Secret clearance.
C++, SIMD, GPU, HLA, DIS, AFSIM