5 gpu inference performance engineer jobs at 5 companies in Boston, MA

1mo
Save
Mark Applied
Hide
Staff Engineer, Inference Optimizations
Boston, Massachusetts, United States
$191k-$239k/yr RemoteFull Time
DigitalOcean
DigitalOceanNew York Stock Exchange: DOCN: Simplifies cloud infrastructure for developers, startups, and SMBs.
5+ YOE5+ years in high-performance computing or AI infrastructure, deep GPU and inference optimization expertise, experience with CUDA/Triton/ROCm, attention-layer and kernel-level optimization, strong system design and leadership through influence.
AITER, CUDA, ROCm, TensorRT, Triton, FlashAttention
4w
Save
Mark Applied
Hide
Staff/Principal DevOps Engineer, AI Inference
Cambridge, Massachusetts, United States
$192k-$272k/yr OnsiteFull Time
Lila Sciences
Lila Sciences: Develops an AI platform for autonomous scientific research and discovery.
Expertise operating GPU/accelerator infrastructure for ML, Kubernetes and AWS deployment experience, infrastructure-as-code (Terraform, Helm), Python proficiency, networking and performance optimization for low-latency inference.
Kubernetes, vLLM, Triton Inference Server, TGI, Terraform, Helm, EKS, EC2, S3, EFA, IAM, NCCL, Python, Rust, Go, CUDA
1w
Save
Mark Applied
Hide
Member of Technical Staff, Performance & Capacity
Boston, Massachusetts, United States
HybridFull Time
Physical Superintelligence
Physical Superintelligence: Building AI systems to discover new physics at scale.
5+ YOERequires 5+ years with GPU and large-scale multi-node compute workloads, distributed training performance, networking, parallel file systems, AI training and inference optimization, and capacity decisions.
GPU, H100, B200, InfiniBand, RDMA, parallel file systems
2mo
Save
Mark Applied
Hide
Senior Deep Learning Framework Communications Engineer
Santa Clara or Austin or Westford or Durham or United States
$152k-$288k/yr HybridFull Time
NVIDIA
NVIDIANASDAQ: NVDA: Designs graphics processing units and artificial intelligence hardware.
5+ YOE5+ years software engineering experience in HPC/AI, experience with PyTorch/JAX and inference engines, Python/C++/CUDA development, performance benchmarking and profilers, understanding of multi-GPU communication and compilers.
PyTorch, TRT-LLM, vLLM, SGLang, JAX, NCCL, NVSHMEM, GPUDirect, torch.compile, PyTorch profiler, NVIDIA Nsight Systems, Python, C++, CUDA, Triton, cuTe, MPI
3mo
Save
Mark Applied
Hide
Software Solutions Architect
Austin or Boxborough or Markham
$212k-$318k/yr OnsiteFull Time
AMD
AMDNasdaq: AMD: Designs and sells microprocessors and graphics hardware for computers.
Architect and deliver enterprise software solutions leveraging AMD GPUs/APUs; engage customers and partners; strong software engineering, AI inference stack and performance optimization experience.
ROCm, ONNX Runtime, ONNX, PyTorch, TensorRT, CUDA, Docker, Kubernetes, C, C++, Python