20 gpu performance profiling engineer jobs at 16 companies in Vallejo, CA

1mo
Save
Mark Applied
Hide
GPU Kernel Engineer
San Francisco, California, United States
$180k-$280k/yr OnsiteFull Time
TypeSafe AI
TypeSafe AI: Private frontier AI lab building reliable, general AI systems for real-world automation and decision-making.
Deep CUDA/GPU kernel expertise, experience building and optimizing training and inference kernels, LLM training experience, profiling and eliminating performance bottlenecks.
CUDA, CuTe DSL
2mo
Save
Mark Applied
Hide
Member of Technical Staff, TPU & AMD GPU Performance Engineering
San Francisco, California, United States
$200k-$400k/yr OnsiteFull Time
Inferact
Inferact: Private AI infrastructure building open-source software that makes LLM inference faster and cheaper.
Bachelor's or equivalent experience in CS/engineering; hands-on AMD GPU/TPU optimization; experience with ROCm/HIP/Triton, XLA/JAX/Pallas; strong profiling, benchmarking, and cross-platform performance skills.
vLLM, ROCm, HIP, Triton, CK, AITER, TPU, XLA, JAX, Pallas, SGLang, TensorRT-LLM, ATOM, MLIR, LLVM, PyTorch
3mo
Save
Mark Applied
Hide
Performance Engineer
Palo Alto, California, United States
OnsiteFull Time
RadixArk
RadixArk: Private AI infrastructure building open training, inference, and post-training systems for developers and research labs.
Strong systems engineering in performance-critical software; GPU/distributed systems; profiling tools; Python and C++; CUDA/Triton/ROCm/XLA familiarity; LLM inference concepts; ability to debug across software, hardware, and infra layers; strong communication.
CUDA, Triton, Pallas, ROCm, XLA, Python, C++
3mo
Save
Mark Applied
Hide
Member of Technical Staff, Hardware, Performance Engineer
Palo Alto or Austin
$200k-$420k/yr OnsiteFull Time
River AI
River AI: Full-stack AI building personal AI, custom models, training infrastructure, and local hardware for developers and enterprises.
5+ YOEBachelor's in EE/CE/CS and 5+ years experience with advanced process nodes; expert in C/C++ or SystemC; familiarity with compilers (LLVM/GCC/XLA), CUDA/Triton, computer architecture, and profiling/trace analysis.
C, C++, SystemC, LLVM, GCC, XLA, CUDA, Triton, QEMU
3mo
Save
Mark Applied
Hide
Software Engineer, ML Performance Optimization
Foster City, California, United States
$192k-$257k/yr OnsiteFull Time
Zoox
Zoox: Autonomous mobility developing a fully electric robotaxi fleet.
4+ YOE4+ years total exp; 2+ years in large-scale model training or inference; PyTorch; GPU-accelerated inference; profiling tools; Python or C++.
PyTorch, TensorRT, NVIDIA Nsight, Python, C++
1mo
Save
Mark Applied
Hide
Research Scientist / Engineer – Performance Optimization
Redwood City, California, United States
OnsiteFull Time
Luma AI
Luma AI: AI is a private creative AI platform generating video and images for creators and teams.
Expert GPU/CPU/accelerator optimization with Triton/CUDA, strong PyTorch and kernel development, profiling tools experience, deep transformer knowledge, and distributed deployment skills.
Triton, CUDA, PyTorch, NVIDIA Nsight, torch profiler, torch.compile, TensorRT, ONNX, XLA
1w
Save
Mark Applied
Hide
Staff Software Engineer - Video Performance - (Bay area only)
San Francisco, California, United States
$251k-$329k/yr OnsiteFull Time
Canva
Canva: Online graphic design and visual communication platform.
Strong C++ proficiency; experience optimizing multithreaded systems, graphics APIs, multimedia, profiling, telemetry, and CPU/GPU architectures. Systems language experience and performance engineering expertise are valued.
C++, Rust, GLSL, HLSL, Metal, Vulkan, WebGPU, OpenGL, Perf, Instruments, Chrome DevTools, Systrace, H.264, H.265, VP9, AV1, iOS, Android, Web, SIMD
1mo
Save
Mark Applied
Hide
Software Engineer, Runtime
Palo Alto, California, United States
OnsiteFull Time
Ollama
Ollama: Open-model software platform helping developers run models locally and in the cloud.
Experience with systems programming (Go,C,C++), GPU or low-level performance work, profiling and optimizing real workloads, and shipping software across macOS, Linux, and Windows.
Go, C, C++, CUDA, Metal, SYCL, MLX
1w
Save
Mark Applied
Hide
Staff Software Engineer - Video Performance - (Bay area only)
San Francisco, California, United States
$251k-$329k/yr OnsiteFull Time
Canva
Canva: Online graphic design and visual communication platform.
Strong C++ proficiency; systems performance optimization, CPU/GPU architecture, SIMD, graphics APIs, multimedia codecs, profiling, diagnostics, telemetry, and cross-team technical collaboration experience.
C++, Rust, GLSL, HLSL, Metal, Vulkan, WebGPU, OpenGL, Perf, Instruments, Chrome DevTools, Systrace, CMake, Web, iOS, Android, H.264, H.265, VP9, AV1, SIMD
1w
Save
Mark Applied
Hide
Genesis-World: Core Simulation Engine Engineer
San Carlos, California, United States
OnsiteFull Time
Genesis AI
Genesis AI: Private French full-stack robotics building general-purpose robots and physical-AI systems for factories, laboratories, hospitals, and homes.
Strong Python engineering, large-codebase architecture, developer tooling or CI, performance profiling, physics and robotics knowledge, and familiarity with JIT compilation or GPU computing.
Python, CUDA, AMD ROCm, Apple Metal, Vulkan, x86, ARM64, Windows, Linux, macOS, py-spy, Nsight, Xcode Instruments, RenderDoc, Drake, MuJoCo, Newton, Isaac, RaiSim, Jiminy, Brax, PyTorch, Triton, JAX, Numba, Maya, Blender, Houdini, USD, glTF, MJCF, URDF, C++
1w
Save
Mark Applied
Hide
Software Engineer, Hardware Accelerators Performance, GeminiApp, DeepMind
Mountain View or San Francisco
$147k-$210k/yr OnsiteFull Time
Google DeepMind
Google DeepMind: Google DeepMind is a private AI research laboratory developing safe artificial intelligence systems for Alphabet.
2+ YOEBachelor’s degree or equivalent practical experience; 2+ years developing in C++ or Python; profiling tools; hardware accelerator software development; memory-hierarchy or instruction-level tuning.
C++, Python, gProf, Valgrind, VTune, CPUs, GPUs, TPUs, JAX
3mo
Save
Mark Applied
Hide
AI/ML Infrastructure Engineer
San Francisco, California, United States
OnsiteFull Time
Zensors
Zensors: AI technology helping airports, retailers, and other large physical businesses automate operations with spatial intelligence.
BS/MS/PhD in CS or EE; strong C/C++ and Python; model optimization, quantization, pruning; GPU performance tuning; profiling tools; cross-functional collaboration.
C/C++, Python, Nsight Systems, Nsight Compute, PyTorch, CUDA, TensorRT, NVIDIA DeepStream, DALI, FFmpeg, TVM, MLIR, ONNX Runtime, Triton
1mo
Save
Mark Applied
Hide
Software Engineer, Systems ML
Sunnyvale or Menlo Park
$154k-$217k/yr OnsiteFull Time
Meta
MetaNASDAQ: META: Builds technologies that help people connect, find communities, and grow businesses.
6+ YOEBachelor's in CS/CE or equivalent, 6+ years software engineering experience in ML systems or HPC, proficiency with PyTorch/TensorFlow, C++ and Python, distributed systems, and performance profiling.
PyTorch, TensorFlow, C++, Python, CUDA, ROCm, MLIR, LLVM, TVM, XLA, IREE
2mo
Save
Mark Applied
Hide
Staff+ Software Engineer, Inference Runtime
San Francisco or Seattle or New York City
$405k-$485k/yr HybridFull Time
Anthropic
Anthropic: AI research developing safe and steerable AI systems.
Senior IC with deep systems or ML infrastructure experience, hands-on performance profiling and optimization, accelerator ecosystem expertise (CUDA/TPU/Trainium), strong software engineering and cross-org alignment skills, and a relevant bachelor’s degree or equivalent.
Rust, Python, CUDA, XLA, Triton, NeuronX, AWS Neuron, Kubernetes, CI/CD
1mo
Save
Mark Applied
Hide
AI Systems, Model Optimization
Palo Alto or United States
HybridFull Time
Unconventional AI
Unconventional AI: Private AI hardware startup building energy-efficient computing substrates for artificial-intelligence workloads.
MS/PhD (or equivalent) in quantitative field, deep practical experience with ML stack and GPU performance optimization, proficiency in profiling and optimizing large ML codebases.
PyTorch, torch.compile, DDP, FSDP, CUDA, Triton, CUTLASS, Megatron-LM, DeepSpeed
3mo
Save
Mark Applied
Hide
Systems Software Engineer (C++)
San Francisco, California, United States
OnsiteFull Time
Thunder Compute
Thunder Compute: GPU virtualization serving developers, enterprises, and AI innovators with cloud infrastructure.
Expert C++ performance optimization, low-level networking, and production systems ownership.
C++, Low-level networking, GPU optimization, Compilers, Performance profiling
2mo
Save
Mark Applied
Hide
Principal Physics Programmer, C++ (Modeling & Simulation)
San Francisco, California, United States
HybridFull Time
Code Metal
Code Metal: AI software translating, verifying, and optimizing code for mission-critical industries deploying software on specialized hardware.
8+ YOE8+ years professional C++ development building performance-critical real-time systems; experience with game-engine or engine-level systems, data-oriented design, concurrency, profiling, and optimization; eligibility for Secret clearance.
C++, SIMD, GPU, HLA, DIS, AFSIM