29 inference optimization engineer jobs at 15 companies in Washington

4w
Save
Mark Applied
Hide
Senior Inference Engineer, GPU Kernel Optimization
Santa Clara or Austin or New York City or Seattle
$184k-$288k/yr OnsiteFull Time
NVIDIA
NVIDIANASDAQ: NVDA: Designs GPU-accelerated computing and artificial intelligence hardware.
6+ YOE6+ years industry experience; strong Python and C++; hands-on GPU profiling (CUPTI, NSYS, NCU); experience with LLM inference frameworks and GPU kernel optimization; advanced degree or equivalent experience.
Python, C++, CUPTI, NSYS, NCU, TRT-LLM, SGLang, vLLM, CUDA, CUTLASS, Triton, PTX, SASS, LLVM, MLIR, ptxas, FlashInfer
4w
Save
Mark Applied
Hide
Senior Inference Engineer, GPU Kernel Optimization
Santa Clara or Austin or New York City or Seattle
$184k-$288k/yr OnsiteFull Time
NVIDIA
NVIDIANASDAQ: NVDA: Designs graphics processing units and artificial intelligence hardware.
6+ YOEMaster's/PhD or equivalent,6+ years industry experience,agentic AI systems experience,strong Python/C++,GPU profiling (CUPTI,NSYS,NCU),LLM inference frameworks,CUDA/CUTLASS/Triton and PTX/SASS familiarity.
Python, C++, CUPTI, NSYS, NCU, TRT-LLM, SGLang, vLLM, CUDA, CUTLASS, Triton, PTX, SASS, LLVM, MLIR, ptxas
1mo
Save
Mark Applied
Hide
Staff Engineer, Inference Optimizations
Seattle, Washington, United States
$191k-$239k/yr HybridFull Time
DigitalOcean
DigitalOceanNew York Stock Exchange: DOCN: Simplifies cloud infrastructure for developers, startups, and SMBs.
8+ YOE8+ years building/operating multi-tenant distributed systems, 2+ years Go/Golang, 2+ years Kubernetes, strong SRE/observability and production debugging skills, experience optimizing inference and GPU utilization.
Go, Golang, Kubernetes, vLLM, Triton, TensorRT-LLM
2mo
Save
Mark Applied
Hide
Member of Technical Staff — Model Optimization and Inference
Seattle, Washington, United States
$250k-$350k/yr OnsiteFull Time
Nuance Labs
Nuance Labs: A building photorealistic, real-time AI avatars and full-duplex audiovisual systems.
Deep expertise in LLM and diffusion-model inference optimization, KV cache strategies, quantization (INT8/INT4, GPTQ/AWQ), profiling/benchmarking, and strong Python/PyTorch skills; familiarity with CUDA/Triton and inference-serving frameworks.
vLLM, SGLang, TensorRT-LLM, Python, PyTorch, CUDA, Triton
1w
Save
Mark Applied
Hide
Senior Inference Engineer, AGI
Sunnyvale or Boston or Seattle or Los Angeles County
$193k-$262k/yr OnsiteFull Time
Amazon
AmazonNASDAQ: AMZN: Global online retail and cloud computing technology provider.
5+ YOERequires 5+ years software development, 4+ years systems architecture, a computer science bachelor's degree, 2+ years neural inference optimization, GPU optimization, real-time systems, and technical leadership experience.
Nsight Compute, Nsight Systems, vLLM, PyTorch, TensorRT-LLM, CUTLASS, Triton, CUDA, PTX, FlashAttention, NCCL, NVLink, AWS Neuron, Trainium, Microsoft Excel
1mo
Save
Mark Applied
Hide
AI Inference Engineer - Speech
Seattle or San Jose
$152k-$332k/yr HybridFull Time
Zoom
ZoomNasdaq: ZM: Provides a cloud-based platform for video, voice, and collaboration.
3+ YOEMaster's in CS/EE or related,3+ years in speech recognition or model inference,deep learning expertise,experience with Python,C/C++,CUDA,TensorRT,PyTorch,TensorFlow and GPU optimization.
Python, shell, C/C++, PyTorch, TensorFlow, CUDA, TensorRT, CUDA Graphs, NVIDIA GPUs, TPU, BrightHire
1mo
Save
Mark Applied
Hide
Research Engineer - LLM/VLM Inference Optimization (Seed Infra)
Seattle, Washington, United States
OnsiteFull Time
ByteDance
ByteDance: Developing AI-driven content platforms and mobile applications.
Bachelor's in CS/EE/Software, strong C/C++ and Python, experience with PyTorch or TensorFlow, production LLM/VLM inference optimization, GPU familiarity and operator optimization, containerization experience.
C, C++, Python, PyTorch, TensorFlow, CUDA, OpenCL, TensorRT, Triton, CUTLASS, FlashAttention, GEMM, GEMV, Conv2D
2mo
Save
Mark Applied
Hide
Fellow, AI Workload Optimization
Bellevue, Washington, United States
$224k-$384k/yr OnsiteFull Time
AMD
AMDNASDAQ: AMD: Designs and manufactures computer processors and graphics technology.
15+ YOE15+ years software development with 5+ years technical leadership; deep expertise in AI frameworks and ROCm; mastery of performance profiling and distributed training/inference optimization; PhD/Master's or equivalent experience.
PyTorch, JAX, vLLM, SGLang, ROCm, TorchProfiler, ROCm Profiler, Nsight
1w
Save
Mark Applied
Hide
Sr. Machine Learning Engineer, Foundation Models Inference - Cloud OS & Inference
Seattle, Washington, United States
OnsiteFull Time
Apple
AppleNASDAQ: AAPL: Designs and sells consumer electronics, software, and online services.
Build and optimize inference frameworks, services, and tools for large-scale foundation models, supporting low-latency AI services across Apple products.
2mo
Save
Mark Applied
Hide
Staff+ Software Engineer, Inference Runtime
San Francisco or Seattle or New York City
$405k-$485k/yr HybridFull Time
Anthropic
Anthropic: Developing safe and reliable artificial intelligence systems.
Senior IC with deep systems or ML infrastructure experience, hands-on performance profiling and optimization, accelerator ecosystem expertise (CUDA/TPU/Trainium), strong software engineering and cross-org alignment skills, and a relevant bachelor’s degree or equivalent.
Rust, Python, CUDA, XLA, Triton, NeuronX, AWS Neuron, Kubernetes, CI/CD
3mo
Save
Mark Applied
Hide
Staff Software Engineer, Inference
Sunnyvale or Bellevue
$188k-$275k/yr HybridFull Time
CoreWeave
CoreWeaveNASDAQ: CRWV: Cloud platform providing GPU-accelerated infrastructure for AI workloads.
8+ YOE8–12+ years in distributed systems; leadership of cross-team initiatives; proficient in Go, Python or C++; production Kubernetes expertise; low-latency, high-throughput system design and optimization.
Go, Python, C++, Kubernetes, CUDA, NCCL, RDMA, NUMA, TensorRT-LLM, Ray Serve, TorchServe
1mo
Save
Mark Applied
Hide
Machine Learning Engineer
Seattle, Washington, United States
$120k-$180k/yr OnsiteFull Time
Constellation Space
Constellation Space: AI-powered operating system for satellite network management.
BS/MS in CS or Engineering (or equivalent), proven experience deploying ML models to production, strong software engineering in Python and C++, experience with Docker, cloud platforms and MLOps, and optimizing low-latency inference.
Python, C++, Docker
3w
Save
Mark Applied
Hide
Staff Software Engineer - Data Cloud Applied ML
San Francisco or Seattle or New York City
$189k-$315k/yr OnsiteFull Time
Rippling
Rippling: Unified platform managing workforce HR, IT, and finance operations
8+ YOE8+ years software engineering experience, distributed systems ownership, experience training/deploying LLMs, model inference optimization, backend skills in Python/Go/Java, and cloud-native infrastructure (Kubernetes).
Python, Go, Java, Kubernetes
3w
Save
Mark Applied
Hide
Senior Machine Learning Engineer
Chicago or New York City or San Francisco or Seattle or Sunnyvale
$182k-$202k/yr OnsiteFull Time
Uber
UberNYSE: UBER: A technology platform for transportation, delivery, and freight.
4+ YOE4+ years building ML models; BS in CS/CE or related; experience with PyTorch, causal inference or constrained optimization preferred; product and marketplace experience a plus.
PyTorch
1mo
Save
Mark Applied
Hide
Principal Scientist - Data Pipeline Engineer
San Jose or Seattle or San Francisco
$206k-$388k/yr OnsiteFull Time
Adobe
AdobeNASDAQ: ADBE: Provides software for digital media creation and marketing analytics
10+ YOE10+ years in data engineering/ML infrastructure, distributed systems expertise, Python and a systems language, experience with Ray or Spark, GPU inference optimization, large-scale databases and data curation for model training.
Ray, Spark, Python, C++, Rust, Go, Java
2mo
Save
Mark Applied
Hide
Principal Software Engineering - AI Frameworks
Redmond or Mountain View or United States
$143k-$331k/yr HybridFull Time
Microsoft
MicrosoftNASDAQ: MSFT: Develops software, services, devices, and cloud computing solutions.
6+ YOEBachelor's in CS or related + 6+ years engineering experience (or equivalent); coding experience in C, C++, C#, Java, JavaScript, or Python; experience with inference stacks, performance optimization, open-source code; ability to pass Microsoft security screening.
C, C++, C#, Java, JavaScript, Python, ONNX, ONNX Runtime, Foundry Local, VSCode, SQL Server, CLI, SDK, REST API