19 gpu ai kernel development engineer jobs at 7 companies in California

2mo
Save
Mark Applied
Hide
Principal GPU AI Kernel Development Engineer
San Diego, California, United States
$202k-$304k/yr OnsiteFull Time
Qualcomm Technologies, Inc.
Qualcomm Technologies, Inc.: Developing semiconductor, wireless, connectivity, automotive, AI, and computing technologies for device and enterprise customers.
6+ YOEDegree in Computer Engineering/Computer Science/Electrical Engineering (BS/MS/PhD) with 6+ years (PhD) to 8+ years (BS) of relevant engineering experience; GPU experience, technical leadership experience and interaction with senior leadership preferred.
1mo
Save
Mark Applied
Hide
Staff Software Development Engineer: GPU, Computer Vision, AI/ML Ops
Santa Clara, California, United States
OnsiteFull Time
AMD
AMDNASDAQ: AMD: Leader in high-performance computing, graphics, and visualization technologies.
Expert in high-performance C++ and GPU programming (HIP/CUDA), experience with LLMs and AI systems, GPU profiling and kernel optimization, and degree in CS/CE/EE.
C++, HIP, CUDA, ROCm, AMD ROCm Profiler, NVIDIA Nsight, CV-CUDA, cuDNN, NCCL, PyTorch, TensorFlow, JAX
2mo
Save
Mark Applied
Hide
Senior AI Software Engineer, Kernel Libraries
Santa Clara or United States
$184k-$288k/yr RemoteFull Time
NVIDIA
NVIDIANASDAQ: NVDA: Computing platform for AI and accelerated graphics.
6+ YOEMasters in CS/EE or equivalent experience; 6+ years in ML/DL systems; strong Python and C/C++; experience with deep learning frameworks, inference engines, runtimes, and GPU kernel development.
Python, C/C++, PyTorch, JAX, TensorFlow, ONNX, vLLM, SGLang, MLC, FlashInfer, Flash Attention, Apache TVM, MLIR, CUDA C/C++, cuTile, Triton
2mo
Save
Mark Applied
Hide
Senior Staff Software Development Engineer- GPU/AI/ML
Santa Clara, California, United States
$179k-$306k/yr HybridFull Time
AMD
AMDNASDAQ: AMD: Leader in high-performance computing, graphics, and visualization technologies.
Expert C++ and GPU kernel experience (HIP/CUDA), hands-on work on ROCm/CUDA libraries and ML frameworks, deep understanding of transformers/LLMs, Bachelor's in CS/CE/EE or equivalent; Master's/PhD preferred.
C++, HIP, CUDA, ROCm, rocBLAS, hipDNN, Composable Kernel, AITemplate, cuBLAS, cuDNN, CUTLASS, Thrust, CUB, NCCL, PyTorch, TensorFlow, JAX, ROCm Profiler, Nsight, RTL, Verilog, SystemVerilog
3w
Save
Mark Applied
Hide
GPU/AI Application System Software Engineer Intern (System Technologies and Engineering) - 2027 Summer
San Jose, California, United States
OnsiteInternship
ByteDance
ByteDance: Global technology specializing in AI-powered content platforms.
Pursuing a bachelor's or master's degree in computer engineering, electrical engineering, computer science, or related fields; requires OS, Linux kernel, architecture, GPU/CPU benchmarking, and Linux systems experience.
Linux, Python, TensorFlow, PyTorch, MPI, NCCL, UCX, NVSHMEM, RDMA, GPU, TPU, Cloud
1mo
Save
Mark Applied
Hide
Senior Software Engineer - GPU Local AI Platforms
Santa Clara or Westford or Austin or Durham or Seattle
$224k-$431k/yr OnsiteFull Time
NVIDIA
NVIDIANASDAQ: NVDA: Computing platform for AI and accelerated graphics.
12+ YOE12+ years software engineering experience in GPU computing or ML systems, strong Python or C++ skills, GPU kernel optimization (CUDA/Triton), container engineering, and LLM inference knowledge.
Python, C++, CUDA, Triton, Docker, OCI, NCCL, RCCL, NVIDIA Container Toolkit
1mo
Save
Mark Applied
Hide
Senior Software Engineer - GPU Local AI Platforms
Santa Clara or Austin or Westford or Durham or Seattle
$224k-$431k/yr OnsiteFull Time
NVIDIA
NVIDIANASDAQ: NVDA: Computing platform for AI and accelerated graphics.
12+ YOE12+ years software engineering experience with GPU computing or ML systems, strong Python or C++ skills, GPU kernel optimization (CUDA/Triton), container engineering, and LLM inference knowledge.
Python, C++, CUDA, Triton, Docker, OCI, NVIDIA Container Toolkit, NCCL, RCCL
1mo
Save
Mark Applied
Hide
Senior Software Engineer - GPU Local AI Platforms
Santa Clara or Austin or Westford or Durham or Seattle
$224k-$431k/yr OnsiteFull Time
NVIDIA
NVIDIANASDAQ: NVDA: Computing platform for AI and accelerated graphics.
12+ YOERequires BS, MS, PhD, or equivalent experience; 12+ years of software engineering; strong Python or C++; GPU kernel optimization; LLM inference, containers, and performance analysis expertise.
Python, C++, CUDA, Triton, Docker, OCI, NVIDIA Container Toolkit, NCCL, RCCL, CI/CD
2mo
Save
Mark Applied
Hide
Senior Software Engineer, AI and DL Kernel Libraries
Santa Clara or Georgia or Texas or Colorado or Washington or California or Oregon or Massachusetts
$184k-$288k/yr RemoteFull Time
NVIDIA
NVIDIANASDAQ: NVDA: Computing platform for AI and accelerated graphics.
6+ YOEMasters (or equivalent experience) in CS/EE, 6+ years ML/DL systems experience, strong Python and C/C++ skills, GPU kernel development experience (CUDA, Triton, cuTile), familiarity with deep learning frameworks and inference runtimes.
PyTorch, JAX, TensorFlow, ONNX, vLLM, SGLang, MLC, Python, C/C++, CUDA C/C++, cuTile, Triton, FlashInfer, Flash Attention, Apache TVM, MLIR
3w
Save
Mark Applied
Hide
Sr. Software Engineer - AI Triton Kernels
San Jose, California, United States
$146k-$218k/yr HybridFull Time
AMD
AMDNASDAQ: AMD: Leader in high-performance computing, graphics, and visualization technologies.
Deep expertise in GPU kernels, Triton, compiler backends, GPU architecture, parallel algorithms, and performance engineering for AI/ML workloads; bachelor's or master's degree preferred.
Triton, Gluon, PyTorch, vLLM, SGLang, ROCm, LLVM, LLVM AMDGPU, MLIR, IREE, Torch
2mo
Save
Mark Applied
Hide
Sr. System Development Engineer, Edge & High Performance Accelerator Servers for AI/ML
Austin or Seattle or Cupertino
$151k-$235k/yr OnsiteFull Time
Amazon
AmazonNASDAQ: AMZN: Multinational technology focused on e-commerce and cloud computing.
6+ YOE6+ years systems/software development and systems design experience; strong programming in C++, C#, Java, Python, Golang, PowerShell, or Ruby; Linux/Unix experience; experience building reliable, scalable automation, diagnostics, and CI/CD for server fleets.
C++, C#, Java, Python, Golang, PowerShell, Ruby, Linux, Linux kernel, CI/CD, BMC/IPMI, PCIe, NVMe, GPU, ARM, x86
2w
Save
Mark Applied
Hide
Software Engineer, AI Kernels & Performance Optimization — MTIA Software
Bellevue or Menlo Park or New York City
$184k-$257k/yr OnsiteFull Time
Meta
MetaNASDAQ: META: Builds technologies that help people connect, find communities, and grow businesses.
6+ YOEBachelor's degree or equivalent practical experience; 6+ years in HPC, accelerator kernels, compiler backends, or systems performance; C++ and Python proficiency; parallel architecture kernel optimization experience.
C++, Python, CUDA, ROCm/HIP, SYCL/OpenCL, NCCL, RCCL, MLIR, LLVM, TVM, XLA, Halide, PyTorch, torch.compile, Inductor, vLLM, SGLang, CUTLASS, cuBLAS, cuDNN, CUTE, Triton, Helion, ThunderKittens, oneDNN, Composable Kernel, FP8, E4M3, E5M2, MX, INT8, INT4, GEMM, DMA
1w
Save
Mark Applied
Hide
Engineer, Staff – Ambient AI Technology
San Diego, California, United States
$135k-$202k/yr OnsiteFull Time
Qualcomm Technologies, Inc.
Qualcomm Technologies, Inc.: Developing semiconductor, wireless, connectivity, automotive, AI, and computing technologies for device and enterprise customers.
4+ YOEBachelor's with 4+ years, master's with 3+ years, or PhD with 2+ years in software engineering; 2+ years using C, C++, Java, or Python. Architecture and embedded systems experience preferred.
C, C++, Java, Python, ARM, Linux, Linux kernel, IPC, CPU, DSP, NPU, GPU, Android, Windows
2mo
Save
Mark Applied
Hide
Integration Engineer - Perception and Platform
San Jose, California, United States
$179k-$306k/yr RemoteFull Time
AMD
AMDNASDAQ: AMD: Leader in high-performance computing, graphics, and visualization technologies.
Senior software engineer for Physical AI: develop and optimize perception, computer vision, and AI workloads for embedded/edge systems. Strong C++, Python, Linux skills; experience deploying ML/DL on edge, performance profiling, and hardware/software integration.
Vitis AI, AMD Versal AI Edge, C++, Python, Linux, ONNX, PyTorch, TensorFlow, ROS 2, OpenGL, Vulkan, Wayland, Linux kernel, Zynq UltraScale+, Kria, AI Engines, CPU, GPU, NPU, FPGA
2mo
Save
Mark Applied
Hide
Senior Linux Software Engineer
Austin or Santa Clara
$147k-$252k/yr OnsiteFull Time
AMD
AMDNASDAQ: AMD: Leader in high-performance computing, graphics, and visualization technologies.
Senior Linux software engineer with OS partner collaboration to enable Linux distros on AMD CPUs/GPUs; strong problem solving and communication.
C++, Python, C, Linux Kernel, Linux, virtualization, AI tools
2mo
Save
Mark Applied
Hide
Principal Software Engineer - Kernels
Santa Clara, California, United States
$200k-$300k/yr HybridFull Time
d-Matrix
d-Matrix: Private AI infrastructure serving data centers with inference accelerators, networking, and software.
12+ YOEMS with 12+ years or PhD with 7+ years; strong computer architecture; C/C++ and Python in Linux; experience with GPUs/AI accelerators; ML workloads; hardware-software co-design; leadership.
C/C++, Python, Linux, CUDA, TensorFlow, PyTorch, MLIR, LLVM, TVM
3mo
Save
Mark Applied
Hide
HPE Labs - Research Engineer
Milpitas, California, United States
$120k-$243k/yr OnsiteFull Time
Hewlett Packard Enterprise
Hewlett Packard EnterpriseNYSE: HPE: Global edge-to-cloud advancing how people live and work.
Ph.D. in Computer Science or Electrical/Computer Engineering; strong distributed systems, HPC, AI workloads; programming in C/C++, Python; Linux kernel, Open vSwitch; CUDA, PyTorch, Kubernetes; network protocols.
C/C++, Python, Linux, CUDA, PyTorch, TensorFlow, Kubernetes, KNative, Open vSwitch, DPDK, eBPF, IP, TCP/UDP, P4, 5G, USRP
1mo
Save
Mark Applied
Hide
Solutions Architect, Agentic Optimization
Santa Clara or United States
$152k-$242k/yr RemoteFull Time
NVIDIA
NVIDIANASDAQ: NVDA: Computing platform for AI and accelerated graphics.
5+ YOEBS, MS, PhD, or equivalent experience; 5+ years in AI/software engineering; Python or C++; GPU model optimization, kernel development, communication, and customer-facing collaboration.
Python, C++, GPUs, GEMM, GitHub, MLOps, containers, Kubernetes
1mo
Save
Mark Applied
Hide
Solutions Architect, Agentic Optimization
Santa Clara or United States
$152k-$242k/yr HybridFull Time
NVIDIA
NVIDIANASDAQ: NVDA: Computing platform for AI and accelerated graphics.
5+ YOEDegree in CS/EE/Physics/Math or equivalent experience,5+ years AI/software engineering,proficient in Python/C++,experience profiling and optimizing models on GPUs and developing GPU kernels,strong communication skills.
Python, C++, GitHub, Kubernetes, containers, GPUs