73 gpu performance profiling engineer jobs at 31 companies in Pleasanton, CA
3w
Save
Mark Applied
Hide
3w
GPU Performance Profiling Engineer
Austin or Santa Clara
$152k-$288k/yrOnsiteFull Time
NVIDIANASDAQ: NVDA: Computing platform for AI and accelerated graphics.
2+ YOEBachelor's in electrical engineering or computer science with 4+ years, master's with 2+ years, or Ph.D.; expertise in C/C++/Python, drivers, GPU architecture, performance analysis, and hardware-software co-design.
NVIDIANASDAQ: NVDA: Computing platform for AI and accelerated graphics.
2+ YOERequires a BS with 4+ years, MS with 2+ years, or PhD; expertise in C/C++/Python software stacks, user- and kernel-mode drivers, CPU/GPU architecture, hardware pipelines, performance analysis, and hardware-software co-design.
NVIDIANASDAQ: NVDA: Computing platform for AI and accelerated graphics.
2+ YOERequires a bachelor's degree and 4+ years, master's degree and 2+ years, or Ph.D.; expertise in C/C++/Python, drivers, CPU/GPU architecture, hardware pipelines, and performance analysis.
Samsung ElectronicsKorea Exchange: 005930: Global leader in technology, semiconductors, and consumer electronics.
2+ YOE2+ years experience with a Bachelor’s in computer science/engineering (or Master’s). Strong GPU architecture and performance verification experience; proficiency in C++, Python, Linux; familiarity with OpenGL/Vulkan/OpenCL; performance test and profiling experience.
TypeSafe AI: Private frontier AI lab building reliable, general AI systems for real-world automation and decision-making.
Deep CUDA/GPU kernel expertise, experience building and optimizing training and inference kernels, LLM training experience, profiling and eliminating performance bottlenecks.
RadixArk: Private AI infrastructure building open training, inference, and post-training systems for developers and research labs.
Strong systems engineering in performance-critical software; GPU/distributed systems; profiling tools; Python and C++; CUDA/Triton/ROCm/XLA familiarity; LLM inference concepts; ability to debug across software, hardware, and infra layers; strong communication.
Efficient Computer: Energy-efficient processor building programmable general-purpose chips for AI and edge applications.
5+ YOE5+ years experience developing low-level C/C++ libraries for performance on hardware; experience with at least two RISC/DSP/GPU platforms, CUDA or HIP, profiling/benchmarking, EM/real-time library design, strong communication, and a Bachelor's in a technical field.
3+ YOEBachelor's or Master's in computer science, computer/electrical engineering, or related field; 3+ years AI performance engineering; Python, CUDA/C++, GPU architecture, profiling, NCCL, and networking expertise.
Samsung SemiconductorKorea Exchange (KRX): 005930: Global leader in semiconductor solutions including memory, system LSI, and foundry services.
0+ YOEAdvanced degree in CS/CE/EE or equivalent experience; strong LLM systems and NVIDIA GPU performance knowledge; experience profiling AI workloads with Nsight tools; proficiency in Python and C++; experience with PyTorch, DeepSpeed, Ray, or similar.
CoreWeaveNasdaq: CRWV: Specialized cloud provider for large-scale AI and machine learning.
5+ YOE5+ years building HPC/GPU software, hands-on CUDA kernel authoring and optimization, C++/Python coding, GPU profiling, and experience delivering performance at scale.
Member of Technical Staff, Hardware, Performance Engineer
Palo Alto or Austin
$200k-$420k/yrOnsiteFull Time
River AI: Full-stack AI building personal AI, custom models, training infrastructure, and local hardware for developers and enterprises.
5+ YOEBachelor's in EE/CE/CS and 5+ years experience with advanced process nodes; expert in C/C++ or SystemC; familiarity with compilers (LLVM/GCC/XLA), CUDA/Triton, computer architecture, and profiling/trace analysis.
Etched: Private semiconductor startup building AI inference chips, racks, and software for frontier-model customers.
Strong C++ or Rust programming, computer architecture and low-level systems knowledge, interest in performance analysis and profiling, familiarity with profiling tools and OS/driver concepts.
Research Scientist / Engineer – Performance Optimization
Redwood City, California, United States
OnsiteFull Time
Luma AI: AI is a private creative AI platform generating video and images for creators and teams.
Expert GPU/CPU/accelerator optimization with Triton/CUDA, strong PyTorch and kernel development, profiling tools experience, deep transformer knowledge, and distributed deployment skills.
Staff Software Engineer - Video Performance - (Bay area only)
San Francisco, California, United States
$251k-$329k/yrOnsiteFull Time
Canva: Online graphic design and visual communication platform.
Strong C++ proficiency; experience optimizing multithreaded systems, graphics APIs, multimedia, profiling, telemetry, and CPU/GPU architectures. Systems language experience and performance engineering expertise are valued.
Ollama: Open-model software platform helping developers run models locally and in the cloud.
Experience with systems programming (Go,C,C++), GPU or low-level performance work, profiling and optimizing real workloads, and shipping software across macOS, Linux, and Windows.
Senior AI Infrastructure Engineer - Model Training
Mountain View, California, United States
$190k-$260k/yrOnsiteFull Time
Kodiak RoboticsNasdaq: KDK: Public autonomous-vehicle technology serving commercial trucking, industrial trucking, defense, and public-sector customers.
2+ YOEDegree in CS or related field,2+ years ML systems experience,expertise in distributed training,high-performance data pipelines,GPU performance and profiling,Python and PyTorch skills.