73 gpu performance profiling engineer jobs at 31 companies in San Francisco, CA

3w
Save
Mark Applied
Hide
GPU Performance Profiling Engineer
Austin or Santa Clara
$152k-$288k/yr OnsiteFull Time
NVIDIA
NVIDIANASDAQ: NVDA: Computing platform for AI and accelerated graphics.
2+ YOEBachelor's in electrical engineering or computer science with 4+ years, master's with 2+ years, or Ph.D.; expertise in C/C++/Python, drivers, GPU architecture, performance analysis, and hardware-software co-design.
C, C++, Python, CUDA, x86, ARM, GPU
3w
Save
Mark Applied
Hide
GPU Performance Profiling Engineer
Austin or Santa Clara
$152k-$288k/yr OnsiteFull Time
NVIDIA
NVIDIANASDAQ: NVDA: Computing platform for AI and accelerated graphics.
2+ YOERequires a BS with 4+ years, MS with 2+ years, or PhD; expertise in C/C++/Python software stacks, user- and kernel-mode drivers, CPU/GPU architecture, hardware pipelines, performance analysis, and hardware-software co-design.
C, C++, Python, CUDA, x86, ARM
3w
Save
Mark Applied
Hide
GPU Performance Profiling Engineer
Austin or Santa Clara
$152k-$288k/yr OnsiteFull Time
NVIDIA
NVIDIANASDAQ: NVDA: Computing platform for AI and accelerated graphics.
2+ YOERequires a bachelor's degree and 4+ years, master's degree and 2+ years, or Ph.D.; expertise in C/C++/Python, drivers, CPU/GPU architecture, hardware pipelines, and performance analysis.
C, C++, Python, CUDA, x86, ARM, GPU
1mo
Save
Mark Applied
Hide
Senior GPU Inference Performance Engineer
Santa Clara, California, United States
$144k-$246k/yr HybridFull Time
AMD
AMDNASDAQ: AMD: Leader in high-performance computing, graphics, and visualization technologies.
Hands-on GPU performance engineering experience; proficent in GPU profiling, distributed inference, AI serving frameworks, Python and C/C++; bachelor's degree preferred.
ROCm, rocProfiler, Omniperf, Omnitrace, CUDA, Nsight Systems, Nsight Compute, DCGM, vLLM, SGLang, FP8, MXFP4, INT4, AWQ, GPTQ, RDMA, RoCE, NCCL, RCCL, Pensando, ConnectX, Docker, containerd, Kubernetes, Python, C++, C, HIP, perftest, ib_write_bw, tcpdump
2mo
Save
Mark Applied
Hide
Senior Engineer, GPU Graphic Core Performance Verification
San Jose, California, United States
$139k-$208k/yr OnsiteFull Time
Samsung Electronics
Samsung ElectronicsKorea Exchange: 005930: Global leader in technology, semiconductors, and consumer electronics.
2+ YOE2+ years experience with a Bachelor’s in computer science/engineering (or Master’s). Strong GPU architecture and performance verification experience; proficiency in C++, Python, Linux; familiarity with OpenGL/Vulkan/OpenCL; performance test and profiling experience.
C++, Python, Linux, OpenGL, Vulkan, OpenCL, SystemC, Sparta, RTL
1mo
Save
Mark Applied
Hide
GPU Kernel Engineer
San Francisco, California, United States
$180k-$280k/yr OnsiteFull Time
TypeSafe AI
TypeSafe AI: Private frontier AI lab building reliable, general AI systems for real-world automation and decision-making.
Deep CUDA/GPU kernel expertise, experience building and optimizing training and inference kernels, LLM training experience, profiling and eliminating performance bottlenecks.
CUDA, CuTe DSL
2mo
Save
Mark Applied
Hide
Member of Technical Staff, TPU & AMD GPU Performance Engineering
San Francisco, California, United States
$200k-$400k/yr OnsiteFull Time
Inferact
Inferact: Private AI infrastructure building open-source software that makes LLM inference faster and cheaper.
Bachelor's or equivalent experience in CS/engineering; hands-on AMD GPU/TPU optimization; experience with ROCm/HIP/Triton, XLA/JAX/Pallas; strong profiling, benchmarking, and cross-platform performance skills.
vLLM, ROCm, HIP, Triton, CK, AITER, TPU, XLA, JAX, Pallas, SGLang, TensorRT-LLM, ATOM, MLIR, LLVM, PyTorch
3mo
Save
Mark Applied
Hide
Performance Engineer
Palo Alto, California, United States
OnsiteFull Time
RadixArk
RadixArk: Private AI infrastructure building open training, inference, and post-training systems for developers and research labs.
Strong systems engineering in performance-critical software; GPU/distributed systems; profiling tools; Python and C++; CUDA/Triton/ROCm/XLA familiarity; LLM inference concepts; ability to debug across software, hardware, and infra layers; strong communication.
CUDA, Triton, Pallas, ROCm, XLA, Python, C++
1mo
Save
Mark Applied
Hide
Performance Library Engineer
San Jose or Pittsburgh
$130k-$230k/yr OnsiteFull Time
Efficient Computer
Efficient Computer: Energy-efficient processor building programmable general-purpose chips for AI and edge applications.
5+ YOE5+ years experience developing low-level C/C++ libraries for performance on hardware; experience with at least two RISC/DSP/GPU platforms, CUDA or HIP, profiling/benchmarking, EM/real-time library design, strong communication, and a Bachelor's in a technical field.
CUDA, HIP, PTX, LLVM IR, MLIR, C, C++
2w
Save
Mark Applied
Hide
Helix AI Engineer, Training Performance
San Jose, California, United States
$200k-$400k/yr OnsiteFull Time
Figure
Figure: Develops general-purpose humanoid robots.
3+ YOEBachelor's or Master's in computer science, computer/electrical engineering, or related field; 3+ years AI performance engineering; Python, CUDA/C++, GPU architecture, profiling, NCCL, and networking expertise.
Triton, CUDA, Nsight Systems, Nsight Compute, PyTorch Profiler, HTA, NCCL, RDMA, NVLink, InfiniBand, RoCE, Python, CUDA/C++, PyTorch, Megatron-LM, vLLM, DeepSpeed, JAX, FSDP, Gluon
1mo
Save
Mark Applied
Hide
Senior Performance Engineer
San Jose, California, United States
$138k-$206k/yr OnsiteFull Time
Samsung Semiconductor
Samsung SemiconductorKorea Exchange (KRX): 005930: Global leader in semiconductor solutions including memory, system LSI, and foundry services.
0+ YOEAdvanced degree in CS/CE/EE or equivalent experience; strong LLM systems and NVIDIA GPU performance knowledge; experience profiling AI workloads with Nsight tools; proficiency in Python and C++; experience with PyTorch, DeepSpeed, Ray, or similar.
Nsight Systems, Nsight Compute, Python, C++, PyTorch, vLLM, SGLang, TensorRT-LLM, DeepSpeed, Ray, Megatron-LM
1mo
Save
Mark Applied
Hide
Senior Software Engineer - GPU Kernel Authoring & Optimization
Sunnyvale or Bellevue
$182k-$242k/yr OnsiteFull Time
CoreWeave
CoreWeaveNasdaq: CRWV: Specialized cloud provider for large-scale AI and machine learning.
5+ YOE5+ years building HPC/GPU software, hands-on CUDA kernel authoring and optimization, C++/Python coding, GPU profiling, and experience delivering performance at scale.
CUDA, Nsight Compute, Nsight Systems, C++, Python, Triton, Mojo, CuTe DSL, JAX, HIP, ROCm, NCCL, Kubernetes, SUNK, Slurm, MLPerf, vLLM, TensorRT-LLM, llm-d, SGLang, KNYFE, Pallas, CUTLASS
3mo
Save
Mark Applied
Hide
Member of Technical Staff, Hardware, Performance Engineer
Palo Alto or Austin
$200k-$420k/yr OnsiteFull Time
River AI
River AI: Full-stack AI building personal AI, custom models, training infrastructure, and local hardware for developers and enterprises.
5+ YOEBachelor's in EE/CE/CS and 5+ years experience with advanced process nodes; expert in C/C++ or SystemC; familiarity with compilers (LLVM/GCC/XLA), CUDA/Triton, computer architecture, and profiling/trace analysis.
C, C++, SystemC, LLVM, GCC, XLA, CUDA, Triton, QEMU
3w
Save
Mark Applied
Hide
Performance Tools Intern
San Jose, California, United States
OnsiteFull Time, Internship
Etched
Etched: Private semiconductor startup building AI inference chips, racks, and software for frontier-model customers.
Strong C++ or Rust programming, computer architecture and low-level systems knowledge, interest in performance analysis and profiling, familiarity with profiling tools and OS/driver concepts.
C++, Rust, Python, Nsight, VTune, Xprof, Perfetto, PCIe, Linux, Windows, GPU, TPU
3mo
Save
Mark Applied
Hide
Software Engineer, ML Performance Optimization
Foster City, California, United States
$192k-$257k/yr OnsiteFull Time
Zoox
Zoox: Autonomous mobility developing a fully electric robotaxi fleet.
4+ YOE4+ years total exp; 2+ years in large-scale model training or inference; PyTorch; GPU-accelerated inference; profiling tools; Python or C++.
PyTorch, TensorRT, NVIDIA Nsight, Python, C++
1mo
Save
Mark Applied
Hide
Research Scientist / Engineer – Performance Optimization
Redwood City, California, United States
OnsiteFull Time
Luma AI
Luma AI: AI is a private creative AI platform generating video and images for creators and teams.
Expert GPU/CPU/accelerator optimization with Triton/CUDA, strong PyTorch and kernel development, profiling tools experience, deep transformer knowledge, and distributed deployment skills.
Triton, CUDA, PyTorch, NVIDIA Nsight, torch profiler, torch.compile, TensorRT, ONNX, XLA
1w
Save
Mark Applied
Hide
Staff Software Engineer - Video Performance - (Bay area only)
San Francisco, California, United States
$251k-$329k/yr OnsiteFull Time
Canva
Canva: Online graphic design and visual communication platform.
Strong C++ proficiency; experience optimizing multithreaded systems, graphics APIs, multimedia, profiling, telemetry, and CPU/GPU architectures. Systems language experience and performance engineering expertise are valued.
C++, Rust, GLSL, HLSL, Metal, Vulkan, WebGPU, OpenGL, Perf, Instruments, Chrome DevTools, Systrace, H.264, H.265, VP9, AV1, iOS, Android, Web, SIMD
1mo
Save
Mark Applied
Hide
Software Engineer, Runtime
Palo Alto, California, United States
OnsiteFull Time
Ollama
Ollama: Open-model software platform helping developers run models locally and in the cloud.
Experience with systems programming (Go,C,C++), GPU or low-level performance work, profiling and optimizing real workloads, and shipping software across macOS, Linux, and Windows.
Go, C, C++, CUDA, Metal, SYCL, MLX
1mo
Save
Mark Applied
Hide
Senior AI Infrastructure Engineer - Model Training
Mountain View, California, United States
$190k-$260k/yr OnsiteFull Time
Kodiak Robotics
Kodiak RoboticsNasdaq: KDK: Public autonomous-vehicle technology serving commercial trucking, industrial trucking, defense, and public-sector customers.
2+ YOEDegree in CS or related field,2+ years ML systems experience,expertise in distributed training,high-performance data pipelines,GPU performance and profiling,Python and PyTorch skills.
PyTorch, PyTorch DDP/FSDP, DeepSpeed, Megatron, NCCL, WebDataset, MosaicML Streaming, MDS, Nsight, PyTorch Profiler, Python, C++, CUDA, Triton, NVLink, InfiniBand
1w
Save
Mark Applied
Hide
Staff Software Engineer - Video Performance - (Bay area only)
San Francisco, California, United States
$251k-$329k/yr OnsiteFull Time
Canva
Canva: Online graphic design and visual communication platform.
Strong C++ proficiency; systems performance optimization, CPU/GPU architecture, SIMD, graphics APIs, multimedia codecs, profiling, diagnostics, telemetry, and cross-team technical collaboration experience.
C++, Rust, GLSL, HLSL, Metal, Vulkan, WebGPU, OpenGL, Perf, Instruments, Chrome DevTools, Systrace, CMake, Web, iOS, Android, H.264, H.265, VP9, AV1, SIMD