5 inference optimization engineer jobs at 2 companies in District of Columbia

2w
Save
Mark Applied
Hide
Machine Learning Performance Engineer - Offboard Training & Inference
Sunnyvale or Washington, D.C. or San Diego or Fort Walton Beach or Ann Arbor or London or Stuttgart or Munich or Stockholm or Bangalore or Seoul or Tokyo
$215k-$285k/yr OnsiteFull Time
Applied Intuition
Applied Intuition: Providing digital infrastructure for physical AI and autonomy.
ML performance engineering experience with distributed training, batch inference, GPU or accelerator optimization, Python, and C++ or another systems language; strong debugging and analytical skills required.
FSDP, DeepSpeed, Megatron, NCCL, NVIDIA Triton Inference Server, TensorRT, ONNX Runtime, Ray, Python, C++, CUDA, Triton, CUTLASS, Nsight Systems, Nsight Compute, PyTorch Profiler, perf, Kubernetes, Slurm, ROS, OpenCV
3w
Save
Mark Applied
Hide
Engineering Manager, Deep Learning Inference
Santa Clara or Washington or Texas or New York or Washington or Massachusetts
$184k-$357k/yr HybridFull Time
NVIDIA
NVIDIANASDAQ: NVDA: Computing platform for AI and accelerated graphics.
6+ YOE3+ MgmtRequires an MS, PhD, or equivalent experience in computer science, engineering, or related field; 6+ years software development; 3+ years technical leadership or engineering management; C/C++, GPU programming, and optimization experience.
vLLM, SGLang, FlashInfer, CUDA, Triton, CUTLASS, NIXL, NCCL, NVSHMEM, C/C++, Python, PyTorch, TensorRT-LLM, Agile
3w
Save
Mark Applied
Hide
Engineering Manager, Deep Learning Inference
Santa Clara or Washington or Texas or New York or Washington or Massachusetts
$184k-$357k/yr RemoteFull Time
NVIDIA
NVIDIANASDAQ: NVDA: Computing platform for AI and accelerated graphics.
6+ YOE3+ MgmtRequires MS, PhD, or equivalent experience; 6+ years software development; 3+ years technical leadership or engineering management; C/C++, GPU programming, performance optimization, and production deep learning deployment.
vLLM, SGLang, FlashInfer, CUDA, Triton, CUTLASS, NIXL, NCCL, NVSHMEM, C/C++, Python, PyTorch, TensorRT-LLM, Agile
1mo
Save
Mark Applied
Hide
Engineering Manager, Deep Learning Inference
Santa Clara or Georgia or District of Columbia or Illinois or California or Massachusetts
$224k-$431k/yr HybridFull Time
NVIDIA
NVIDIANASDAQ: NVDA: Computing platform for AI and accelerated graphics.
6+ YOE3+ MgmtMS/PhD or equivalent, 6+ years software experience with 3+ years technical leadership; strong C/C++ and Python; GPU programming and performance optimization experience; experience with deep learning model deployment.
SGLang, vLLM, FlashInfer, C, C++, Python, CUDA, Triton, CUTLASS, NIXL, NCCL, NVSHMEM, PyTorch, TensorRT-LLM
1mo
Save
Mark Applied
Hide
Engineering Manager, Deep Learning Inference
Santa Clara or Washington or Georgia or Illinois or California or Massachusetts
$224k-$431k/yr HybridFull Time
NVIDIA
NVIDIANASDAQ: NVDA: Computing platform for AI and accelerated graphics.
6+ YOE3+ MgmtMS/PhD or equivalent,6+ years software development with 3+ years technical leadership,strong C/C++ and Python,experience with GPU programming (CUDA,Triton,CUTLASS) and deploying/optimizing deep learning models,Agile experience.
SGLang, vLLM, FlashInfer, C/C++, Python, CUDA, Triton, CUTLASS, NIXL, NCCL, NVSHMEM, PyTorch, TensorRT-LLM