4 gpu ai kernel development engineer jobs at 1 company in Massachusetts

1mo
Save
Mark Applied
Hide
Senior Software Engineer - GPU Local AI Platforms
Santa Clara or Austin or Westford or Durham or Seattle
$224k-$431k/yr OnsiteFull Time
NVIDIA
NVIDIANASDAQ: NVDA: Computing platform for AI and accelerated graphics.
12+ YOE12+ years software engineering experience with GPU computing or ML systems, strong Python or C++ skills, GPU kernel optimization (CUDA/Triton), container engineering, and LLM inference knowledge.
Python, C++, CUDA, Triton, Docker, OCI, NVIDIA Container Toolkit, NCCL, RCCL
1mo
Save
Mark Applied
Hide
Senior Software Engineer - GPU Local AI Platforms
Santa Clara or Westford or Austin or Durham or Seattle
$224k-$431k/yr OnsiteFull Time
NVIDIA
NVIDIANASDAQ: NVDA: Computing platform for AI and accelerated graphics.
12+ YOE12+ years software engineering experience in GPU computing or ML systems, strong Python or C++ skills, GPU kernel optimization (CUDA/Triton), container engineering, and LLM inference knowledge.
Python, C++, CUDA, Triton, Docker, OCI, NCCL, RCCL, NVIDIA Container Toolkit
1mo
Save
Mark Applied
Hide
Senior Software Engineer - GPU Local AI Platforms
Santa Clara or Austin or Westford or Durham or Seattle
$224k-$431k/yr OnsiteFull Time
NVIDIA
NVIDIANASDAQ: NVDA: Computing platform for AI and accelerated graphics.
12+ YOERequires BS, MS, PhD, or equivalent experience; 12+ years of software engineering; strong Python or C++; GPU kernel optimization; LLM inference, containers, and performance analysis expertise.
Python, C++, CUDA, Triton, Docker, OCI, NVIDIA Container Toolkit, NCCL, RCCL, CI/CD
2mo
Save
Mark Applied
Hide
Senior Software Engineer, AI and DL Kernel Libraries
Santa Clara or Georgia or Texas or Colorado or Washington or California or Oregon or Massachusetts
$184k-$288k/yr RemoteFull Time
NVIDIA
NVIDIANASDAQ: NVDA: Computing platform for AI and accelerated graphics.
6+ YOEMasters (or equivalent experience) in CS/EE, 6+ years ML/DL systems experience, strong Python and C/C++ skills, GPU kernel development experience (CUDA, Triton, cuTile), familiarity with deep learning frameworks and inference runtimes.
PyTorch, JAX, TensorFlow, ONNX, vLLM, SGLang, MLC, Python, C/C++, CUDA C/C++, cuTile, Triton, FlashInfer, Flash Attention, Apache TVM, MLIR