22 gpu ai kernel development engineer jobs at 9 companies in United States

2mo
Save
Mark Applied
Hide
Principal GPU AI Kernel Development Engineer
San Diego, California, United States
$202k-$304k/yr OnsiteFull Time
Qualcomm Technologies, Inc.
Qualcomm Technologies, Inc.: Developing semiconductor, wireless, connectivity, automotive, AI, and computing technologies for device and enterprise customers.
6+ YOEDegree in Computer Engineering/Computer Science/Electrical Engineering (BS/MS/PhD) with 6+ years (PhD) to 8+ years (BS) of relevant engineering experience; GPU experience, technical leadership experience and interaction with senior leadership preferred.
1mo
Save
Mark Applied
Hide
Staff Software Development Engineer: GPU, Computer Vision, AI/ML Ops
Santa Clara, California, United States
OnsiteFull Time
AMD
AMDNASDAQ: AMD: Leader in high-performance computing, graphics, and visualization technologies.
Expert in high-performance C++ and GPU programming (HIP/CUDA), experience with LLMs and AI systems, GPU profiling and kernel optimization, and degree in CS/CE/EE.
C++, HIP, CUDA, ROCm, AMD ROCm Profiler, NVIDIA Nsight, CV-CUDA, cuDNN, NCCL, PyTorch, TensorFlow, JAX
2mo
Save
Mark Applied
Hide
Senior AI Software Engineer, Kernel Libraries
Santa Clara or United States
$184k-$288k/yr RemoteFull Time
NVIDIA
NVIDIANASDAQ: NVDA: Computing platform for AI and accelerated graphics.
6+ YOEMasters in CS/EE or equivalent experience; 6+ years in ML/DL systems; strong Python and C/C++; experience with deep learning frameworks, inference engines, runtimes, and GPU kernel development.
Python, C/C++, PyTorch, JAX, TensorFlow, ONNX, vLLM, SGLang, MLC, FlashInfer, Flash Attention, Apache TVM, MLIR, CUDA C/C++, cuTile, Triton
3w
Save
Mark Applied
Hide
GPU/AI Application System Software Engineer Intern (System Technologies and Engineering) - 2027 Summer
San Jose, California, United States
OnsiteInternship
ByteDance
ByteDance: Global technology specializing in AI-powered content platforms.
Pursuing a bachelor's or master's degree in computer engineering, electrical engineering, computer science, or related fields; requires OS, Linux kernel, architecture, GPU/CPU benchmarking, and Linux systems experience.
Linux, Python, TensorFlow, PyTorch, MPI, NCCL, UCX, NVSHMEM, RDMA, GPU, TPU, Cloud
1mo
Save
Mark Applied
Hide
Senior Software Engineer - GPU Local AI Platforms
Santa Clara or Westford or Austin or Durham or Seattle
$224k-$431k/yr OnsiteFull Time
NVIDIA
NVIDIANASDAQ: NVDA: Computing platform for AI and accelerated graphics.
12+ YOE12+ years software engineering experience in GPU computing or ML systems, strong Python or C++ skills, GPU kernel optimization (CUDA/Triton), container engineering, and LLM inference knowledge.
Python, C++, CUDA, Triton, Docker, OCI, NCCL, RCCL, NVIDIA Container Toolkit
1mo
Save
Mark Applied
Hide
Senior Software Engineer - GPU Local AI Platforms
Santa Clara or Austin or Westford or Durham or Seattle
$224k-$431k/yr OnsiteFull Time
NVIDIA
NVIDIANASDAQ: NVDA: Computing platform for AI and accelerated graphics.
12+ YOERequires BS, MS, PhD, or equivalent experience; 12+ years of software engineering; strong Python or C++; GPU kernel optimization; LLM inference, containers, and performance analysis expertise.
Python, C++, CUDA, Triton, Docker, OCI, NVIDIA Container Toolkit, NCCL, RCCL, CI/CD
2mo
Save
Mark Applied
Hide
AI Research Engineer, Inference
London or New York City
$250k-$300k/yr OnsiteFull Time
Hudson River Trading
Hudson River Trading: Private quantitative trading firm providing liquidity across global markets and directly to financial-market clients.
2+ YOEStrong engineering skills and 2+ years building deep learning systems; experience with GPU kernels, PyTorch, JAX, XLA, CUDA Graphs, FPGA, or ASICs preferred.
CUDA, Triton, Pallas, CuTe DSL, PyTorch, JAX, XLA, CUDA Graphs, FPGA, ASICs
5d
Save
Mark Applied
Hide
Staff Applied AI Inference Engineer
Denver, Colorado, United States
$185k-$225k/yr OnsiteFull Time
Crusoe
Crusoe: Vertically integrated AI infrastructure and energy.
Bachelor's, master's, or Ph.D. in a related field; production coding experience in Python or C++; LLM inference optimization, serving frameworks, kernel profiling, GPU, AI/ML pipeline, and communication skills.
Python, C++, vLLM, SGLang, CUDA, Docker, Kubernetes
2mo
Save
Mark Applied
Hide
Sr. System Development Engineer, Edge & High Performance Accelerator Servers for AI/ML
Austin or Seattle or Cupertino
$151k-$235k/yr OnsiteFull Time
Amazon
AmazonNASDAQ: AMZN: Multinational technology focused on e-commerce and cloud computing.
6+ YOE6+ years systems/software development and systems design experience; strong programming in C++, C#, Java, Python, Golang, PowerShell, or Ruby; Linux/Unix experience; experience building reliable, scalable automation, diagnostics, and CI/CD for server fleets.
C++, C#, Java, Python, Golang, PowerShell, Ruby, Linux, Linux kernel, CI/CD, BMC/IPMI, PCIe, NVMe, GPU, ARM, x86
2w
Save
Mark Applied
Hide
Software Engineer, AI Kernels & Performance Optimization — MTIA Software
Bellevue or Menlo Park or New York City
$184k-$257k/yr OnsiteFull Time
Meta
MetaNASDAQ: META: Builds technologies that help people connect, find communities, and grow businesses.
6+ YOEBachelor's degree or equivalent practical experience; 6+ years in HPC, accelerator kernels, compiler backends, or systems performance; C++ and Python proficiency; parallel architecture kernel optimization experience.
C++, Python, CUDA, ROCm/HIP, SYCL/OpenCL, NCCL, RCCL, MLIR, LLVM, TVM, XLA, Halide, PyTorch, torch.compile, Inductor, vLLM, SGLang, CUTLASS, cuBLAS, cuDNN, CUTE, Triton, Helion, ThunderKittens, oneDNN, Composable Kernel, FP8, E4M3, E5M2, MX, INT8, INT4, GEMM, DMA
2mo
Save
Mark Applied
Hide
Principal Software Engineer - Kernels
Santa Clara, California, United States
$200k-$300k/yr HybridFull Time
d-Matrix
d-Matrix: Private AI infrastructure serving data centers with inference accelerators, networking, and software.
12+ YOEMS with 12+ years or PhD with 7+ years; strong computer architecture; C/C++ and Python in Linux; experience with GPUs/AI accelerators; ML workloads; hardware-software co-design; leadership.
C/C++, Python, Linux, CUDA, TensorFlow, PyTorch, MLIR, LLVM, TVM
3mo
Save
Mark Applied
Hide
HPE Labs - Research Engineer
Milpitas, California, United States
$120k-$243k/yr OnsiteFull Time
Hewlett Packard Enterprise
Hewlett Packard EnterpriseNYSE: HPE: Global edge-to-cloud advancing how people live and work.
Ph.D. in Computer Science or Electrical/Computer Engineering; strong distributed systems, HPC, AI workloads; programming in C/C++, Python; Linux kernel, Open vSwitch; CUDA, PyTorch, Kubernetes; network protocols.
C/C++, Python, Linux, CUDA, PyTorch, TensorFlow, Kubernetes, KNative, Open vSwitch, DPDK, eBPF, IP, TCP/UDP, P4, 5G, USRP