8 gpu inference performance engineer jobs at 5 companies in Massachusetts
1mo
Save
Mark Applied
Hide
1mo
Staff Engineer, Inference Optimizations
Boston, Massachusetts, United States
$191k-$239k/yrRemoteFull Time
DigitalOceanNew York Stock Exchange: DOCN: Simplifies cloud infrastructure for developers, startups, and SMBs.
5+ YOE5+ years in high-performance computing or AI infrastructure, deep GPU and inference optimization expertise, experience with CUDA/Triton/ROCm, attention-layer and kernel-level optimization, strong system design and leadership through influence.
Santa Clara or Washington or Texas or New York or Washington or Massachusetts
$184k-$357k/yrRemoteFull Time
NVIDIANASDAQ: NVDA: Designs graphics processing units and artificial intelligence hardware.
6+ YOE3+ MgmtRequires MS, PhD, or equivalent experience; 6+ years software development; 3+ years technical leadership or engineering management; C/C++, GPU programming, performance optimization, and production deep learning deployment.
Santa Clara or Georgia or District of Columbia or Illinois or California or Massachusetts
$224k-$431k/yrHybridFull Time
NVIDIANASDAQ: NVDA: Designs GPU-accelerated computing and artificial intelligence hardware.
6+ YOE3+ MgmtMS/PhD or equivalent, 6+ years software experience with 3+ years technical leadership; strong C/C++ and Python; GPU programming and performance optimization experience; experience with deep learning model deployment.
Physical Superintelligence: Building AI systems to discover new physics at scale.
5+ YOERequires 5+ years with GPU and large-scale multi-node compute workloads, distributed training performance, networking, parallel file systems, AI training and inference optimization, and capacity decisions.
GPU, H100, B200, InfiniBand, RDMA, parallel file systems
Senior Deep Learning Framework Communications Engineer
Santa Clara or Austin or Westford or Durham or United States
$152k-$288k/yrHybridFull Time
NVIDIANASDAQ: NVDA: Designs graphics processing units and artificial intelligence hardware.
5+ YOE5+ years software engineering experience in HPC/AI, experience with PyTorch/JAX and inference engines, Python/C++/CUDA development, performance benchmarking and profilers, understanding of multi-GPU communication and compilers.