4 gpu inference performance engineer jobs at 3 companies in Washington

1mo
Save
Mark Applied
Hide
Staff Engineer, Inference Optimizations
Austin or Seattle
$191k-$239k/yr RemoteFull Time
DigitalOcean
DigitalOceanNew York Stock Exchange: DOCN: Simplifies cloud infrastructure for developers, startups, and SMBs.
5+ YOE5+ years in high-performance computing or AI infrastructure with GPU architecture expertise, experience optimizing attention layers and distributed GPU kernels, and strong low-level systems design and open-source contributions.
AITER, CUDA, ROCm, TensorRT, OpenAI Triton
1mo
Save
Mark Applied
Hide
Senior Deep Learning Software Engineer, Inference
California or Texas or New York or Washington or Massachusetts
$152k-$288k/yr RemoteFull Time
NVIDIA
NVIDIANASDAQ: NVDA: Designs GPU-accelerated computing and artificial intelligence hardware.
5+ YOEMasters/PhD or equivalent,5+ years software development,excellent C/C++ skills,CUDA and GPU programming experience preferred,experience optimizing/deploying DL inference,Python and performance profiling experience helpful.
CUTLASS, OAI Triton, NCCL, CUDA, vLLM, SGLang, FlashInfer, PyTorch, NVSHMEM, C/C++, Python
2w
Save
Mark Applied
Hide
Engineering Manager, Deep Learning Inference
Santa Clara or Washington or Texas or New York or Washington or Massachusetts
$184k-$357k/yr RemoteFull Time
NVIDIA
NVIDIANASDAQ: NVDA: Designs graphics processing units and artificial intelligence hardware.
6+ YOE3+ MgmtRequires MS, PhD, or equivalent experience; 6+ years software development; 3+ years technical leadership or engineering management; C/C++, GPU programming, performance optimization, and production deep learning deployment.
vLLM, SGLang, FlashInfer, CUDA, Triton, CUTLASS, NIXL, NCCL, NVSHMEM, C/C++, Python, PyTorch, TensorRT-LLM, Agile
2mo
Save
Mark Applied
Hide
Staff+ Software Engineer, Inference Runtime
San Francisco or Seattle or New York City
$405k-$485k/yr HybridFull Time
Anthropic
Anthropic: Developing safe and reliable artificial intelligence systems.
Senior IC with deep systems or ML infrastructure experience, hands-on performance profiling and optimization, accelerator ecosystem expertise (CUDA/TPU/Trainium), strong software engineering and cross-org alignment skills, and a relevant bachelor’s degree or equivalent.
Rust, Python, CUDA, XLA, Triton, NeuronX, AWS Neuron, Kubernetes, CI/CD