62 gpu kernel development engineer jobs at 21 companies in Concord, CA
4w
Save
Mark Applied
Hide
4w
GPU Kernel Engineer
San Francisco, California, United States
$180k-$280k/yrOnsiteFull Time
TypeSafe AI: Building reliable, general frontier AI models for automation.
Deep CUDA/GPU kernel expertise, experience building and optimizing training and inference kernels, LLM training experience, profiling and eliminating performance bottlenecks.
MakerMaker: Autonomous research agents for recursive self-improvement
4+ YOE4+ years GPU kernel development; hardware-level fluency; profiling; Python/C++ reading; track record of kernel-level optimizations; strong systems know-how.
SR. Software Development Engineer – GPU Kernel Development
Santa Clara, California, United States
$240k-$360k/yrOnsiteFull Time
AMDNASDAQ: AMD: Designs and manufactures computer processors and graphics technology.
Expert C++ and Python developer experienced with GPU kernel development (HIP, CUDA, ASM), deep learning frameworks (TensorFlow, PyTorch), compiler internals (LLVM/ROCm), and performance optimization in Linux; advanced degree preferred.
Bolt Graphics: Designing high-efficiency graphics processors for professional rendering and simulation.
Experience designing and implementing high-performance user/kernel drivers for Windows and Linux; proficiency in C/C++; knowledge of modern GPU APIs and compiler toolchains; BS/MS in related field preferred.
Waymo: Autonomous driving technology for ride-hailing and logistics.
Experience developing GPU kernels and runtimes for machine learning systems; strong ML systems background and ability to work on autonomous driving software.
Principal Engineer, CUDA UMD - GPU Kernel Scheduling
Santa Clara, California, United States
$272k-$431k/yrOnsiteFull Time
NVIDIANASDAQ: NVDA: Designs graphics processing units and artificial intelligence hardware.
15+ YOEBS/MS in CS/EE or equivalent, strong C/C++ skills, 15+ years development experience, OS and multithreading expertise, experience with large codebases and cross-team project leadership; CUDA and kernel-mode experience preferred.
Principal Engineer, CUDA UMD - GPU Kernel Scheduling
Santa Clara, California, United States
$272k-$431k/yrOnsiteFull Time
NVIDIANASDAQ: NVDA: Designs GPU-accelerated computing and artificial intelligence hardware.
15+ YOEBS/MS in CS/EE or equivalent experience, strong C/C++ skills, 15+ years development experience, OS interfaces and multithreading knowledge, experience with large codebases and driving cross-team projects; CUDA and kernel-mode experience preferred.
GPU/AI Application System Software Engineer Intern (System Technologies and Engineering) - 2027 Summer
San Jose, California, United States
OnsiteInternship
ByteDance: Developing AI-driven content platforms and mobile applications.
Pursuing a bachelor's or master's degree in computer engineering, electrical engineering, computer science, or related fields; requires OS, Linux kernel, architecture, GPU/CPU benchmarking, and Linux systems experience.
8+ YOE5+ MgmtBachelor's degree or equivalent practical experience; 8 years software development, 7 years embedded operating systems, 5 years design and architecture, kernel and firmware, and C or C++ experience.
Sr. System Development Engineer, Edge & High Performance Accelerator Servers for AI/ML
Austin or Seattle or Cupertino
$151k-$235k/yrOnsiteFull Time
AmazonNASDAQ: AMZN: Global online retail and cloud computing technology provider.
6+ YOE6+ years systems/software development and systems design experience; strong programming in C++, C#, Java, Python, Golang, PowerShell, or Ruby; Linux/Unix experience; experience building reliable, scalable automation, diagnostics, and CI/CD for server fleets.
C++, C#, Java, Python, Golang, PowerShell, Ruby, Linux, Linux kernel, CI/CD, BMC/IPMI, PCIe, NVMe, GPU, ARM, x86
Silicon Validation Software Engineer- GPU IP Validation and Integration
Cupertino, California, United States
OnsiteFull Time
AppleNASDAQ: AAPL: Designs and sells consumer electronics, software, and online services.
Experience developing graphics/SoC validation software and integrating it into system-level test environments; strong background in graphics, video processing, kernel/embedded systems, and critical attention to detail.
CelesticaNYSE: CLS: Provides design, manufacturing, and supply chain solutions for electronics.
6+ YOEBachelors in Engineering, 6+ years C/C++ and embedded systems experience, hardware bring-up, low-level driver development, Linux kernel and device driver expertise, strong analytical skills.
Research Scientist / Engineer – Performance Optimization
Redwood City, California, United States
OnsiteFull Time
Luma AI: Develops multimodal AI for video generation and creative production.
Expert GPU/CPU/accelerator optimization with Triton/CUDA, strong PyTorch and kernel development, profiling tools experience, deep transformer knowledge, and distributed deployment skills.
Nimble: Building autonomous robotic systems for supply chain fulfillment.
3+ YOE3+ years professional software engineering experience; proficiency in Rust, C, or C++; Linux kernel development, board bring-up, bootloader/uboot/EDK2/UEFI experience, networking knowledge, and strong communication and mentoring skills.
Crusoe: Provides energy-efficient cloud infrastructure powered by stranded and renewable energy.
12+ YOE12+ years designing core infrastructure at hyperscale or in HPC, with Linux kernel, virtualization, high-performance networking, GPU optimization, R&D leadership, and a bachelor's or master's degree in a related field.
Linux kernel, KVM, QEMU, Firecracker, RoCE v2, InfiniBand, RDMA, SR-IOV, Kubernetes, Slurm, NVIDIA, AMD, Linux, IETF RFCs
Nuro: Builds autonomous driving software and electric delivery robots.
2+ YOE2+ years industry experience in robotics or autonomous systems; proficiency in C++ and embedded/real-time systems; familiarity with Linux kernel and device drivers; degree in CS/EE or related; strong communication and collaboration skills.
Black Forest Labs: Develops generative AI models for image and video creation.
Deep experience with large-scale training systems, PyTorch, distributed training, GPU profiling and performance, low-precision/quantization work, kernel-level optimizations, and debugging distributed training failures.
PolyExplore: Develops high-precision mobile mapping and navigation technology.
BA/BS in CS or EE (MA/MS preferred). Strong C/C++ and embedded firmware skills, GPU programming, Linux knowledge, cross-platform development (Windows,iOS,Linux,AWS).
d-Matrix: Develops high-performance semiconductor chips for generative AI inference.
12+ YOEMS with 12+ years or PhD with 7+ years; strong computer architecture; C/C++ and Python in Linux; experience with GPUs/AI accelerators; ML workloads; hardware-software co-design; leadership.