8 gpu inference performance engineer jobs at 5 companies in Massachusetts

1mo
Save
Mark Applied
Hide
Staff Engineer, Inference Optimizations
Boston, Massachusetts, United States
$191k-$239k/yr RemoteFull Time
DigitalOcean
DigitalOceanNew York Stock Exchange: DOCN: Simplifies cloud infrastructure for developers, startups, and SMBs.
5+ YOE5+ years in high-performance computing or AI infrastructure, deep GPU and inference optimization expertise, experience with CUDA/Triton/ROCm, attention-layer and kernel-level optimization, strong system design and leadership through influence.
AITER, CUDA, ROCm, TensorRT, Triton, FlashAttention
4w
Save
Mark Applied
Hide
Staff/Principal DevOps Engineer, AI Inference
Cambridge, Massachusetts, United States
$192k-$272k/yr OnsiteFull Time
Lila Sciences
Lila Sciences: Develops an AI platform for autonomous scientific research and discovery.
Expertise operating GPU/accelerator infrastructure for ML, Kubernetes and AWS deployment experience, infrastructure-as-code (Terraform, Helm), Python proficiency, networking and performance optimization for low-latency inference.
Kubernetes, vLLM, Triton Inference Server, TGI, Terraform, Helm, EKS, EC2, S3, EFA, IAM, NCCL, Python, Rust, Go, CUDA
1mo
Save
Mark Applied
Hide
Senior Deep Learning Software Engineer, Inference
California or Texas or New York or Washington or Massachusetts
$152k-$288k/yr RemoteFull Time
NVIDIA
NVIDIANASDAQ: NVDA: Designs GPU-accelerated computing and artificial intelligence hardware.
5+ YOEMasters/PhD or equivalent,5+ years software development,excellent C/C++ skills,CUDA and GPU programming experience preferred,experience optimizing/deploying DL inference,Python and performance profiling experience helpful.
CUTLASS, OAI Triton, NCCL, CUDA, vLLM, SGLang, FlashInfer, PyTorch, NVSHMEM, C/C++, Python
3w
Save
Mark Applied
Hide
Engineering Manager, Deep Learning Inference
Santa Clara or Washington or Texas or New York or Washington or Massachusetts
$184k-$357k/yr RemoteFull Time
NVIDIA
NVIDIANASDAQ: NVDA: Designs graphics processing units and artificial intelligence hardware.
6+ YOE3+ MgmtRequires MS, PhD, or equivalent experience; 6+ years software development; 3+ years technical leadership or engineering management; C/C++, GPU programming, performance optimization, and production deep learning deployment.
vLLM, SGLang, FlashInfer, CUDA, Triton, CUTLASS, NIXL, NCCL, NVSHMEM, C/C++, Python, PyTorch, TensorRT-LLM, Agile
4w
Save
Mark Applied
Hide
Engineering Manager, Deep Learning Inference
Santa Clara or Georgia or District of Columbia or Illinois or California or Massachusetts
$224k-$431k/yr HybridFull Time
NVIDIA
NVIDIANASDAQ: NVDA: Designs GPU-accelerated computing and artificial intelligence hardware.
6+ YOE3+ MgmtMS/PhD or equivalent, 6+ years software experience with 3+ years technical leadership; strong C/C++ and Python; GPU programming and performance optimization experience; experience with deep learning model deployment.
SGLang, vLLM, FlashInfer, C, C++, Python, CUDA, Triton, CUTLASS, NIXL, NCCL, NVSHMEM, PyTorch, TensorRT-LLM
1w
Save
Mark Applied
Hide
Member of Technical Staff, Performance & Capacity
Boston, Massachusetts, United States
HybridFull Time
Physical Superintelligence
Physical Superintelligence: Building AI systems to discover new physics at scale.
5+ YOERequires 5+ years with GPU and large-scale multi-node compute workloads, distributed training performance, networking, parallel file systems, AI training and inference optimization, and capacity decisions.
GPU, H100, B200, InfiniBand, RDMA, parallel file systems
2mo
Save
Mark Applied
Hide
Senior Deep Learning Framework Communications Engineer
Santa Clara or Austin or Westford or Durham or United States
$152k-$288k/yr HybridFull Time
NVIDIA
NVIDIANASDAQ: NVDA: Designs graphics processing units and artificial intelligence hardware.
5+ YOE5+ years software engineering experience in HPC/AI, experience with PyTorch/JAX and inference engines, Python/C++/CUDA development, performance benchmarking and profilers, understanding of multi-GPU communication and compilers.
PyTorch, TRT-LLM, vLLM, SGLang, JAX, NCCL, NVSHMEM, GPUDirect, torch.compile, PyTorch profiler, NVIDIA Nsight Systems, Python, C++, CUDA, Triton, cuTe, MPI
3mo
Save
Mark Applied
Hide
Software Solutions Architect
Austin or Boxborough or Markham
$212k-$318k/yr OnsiteFull Time
AMD
AMDNasdaq: AMD: Designs and sells microprocessors and graphics hardware for computers.
Architect and deliver enterprise software solutions leveraging AMD GPUs/APUs; engage customers and partners; strong software engineering, AI inference stack and performance optimization experience.
ROCm, ONNX Runtime, ONNX, PyTorch, TensorRT, CUDA, Docker, Kubernetes, C, C++, Python