26 inference optimization engineer jobs at 15 companies in Kirkland, WA

4w
Save
Mark Applied
Hide
Senior Inference Engineer, GPU Kernel Optimization
Santa Clara or Austin or New York City or Seattle
$184k-$288k/yr OnsiteFull Time
NVIDIA
NVIDIANASDAQ: NVDA: Designs GPU-accelerated computing and artificial intelligence hardware.
6+ YOE6+ years industry experience; strong Python and C++; hands-on GPU profiling (CUPTI, NSYS, NCU); experience with LLM inference frameworks and GPU kernel optimization; advanced degree or equivalent experience.
Python, C++, CUPTI, NSYS, NCU, TRT-LLM, SGLang, vLLM, CUDA, CUTLASS, Triton, PTX, SASS, LLVM, MLIR, ptxas, FlashInfer
4w
Save
Mark Applied
Hide
Senior Inference Engineer, GPU Kernel Optimization
Santa Clara or Austin or New York City or Seattle
$184k-$288k/yr OnsiteFull Time
NVIDIA
NVIDIANASDAQ: NVDA: Designs graphics processing units and artificial intelligence hardware.
6+ YOEMaster's/PhD or equivalent,6+ years industry experience,agentic AI systems experience,strong Python/C++,GPU profiling (CUPTI,NSYS,NCU),LLM inference frameworks,CUDA/CUTLASS/Triton and PTX/SASS familiarity.
Python, C++, CUPTI, NSYS, NCU, TRT-LLM, SGLang, vLLM, CUDA, CUTLASS, Triton, PTX, SASS, LLVM, MLIR, ptxas
1mo
Save
Mark Applied
Hide
Staff Engineer, Inference Optimizations
Seattle, Washington, United States
$191k-$239k/yr HybridFull Time
DigitalOcean
DigitalOceanNew York Stock Exchange: DOCN: Simplifies cloud infrastructure for developers, startups, and SMBs.
8+ YOE8+ years building/operating multi-tenant distributed systems, 2+ years Go/Golang, 2+ years Kubernetes, strong SRE/observability and production debugging skills, experience optimizing inference and GPU utilization.
Go, Golang, Kubernetes, vLLM, Triton, TensorRT-LLM
2mo
Save
Mark Applied
Hide
Member of Technical Staff — Model Optimization and Inference
Seattle, Washington, United States
$250k-$350k/yr OnsiteFull Time
Nuance Labs
Nuance Labs: A building photorealistic, real-time AI avatars and full-duplex audiovisual systems.
Deep expertise in LLM and diffusion-model inference optimization, KV cache strategies, quantization (INT8/INT4, GPTQ/AWQ), profiling/benchmarking, and strong Python/PyTorch skills; familiarity with CUDA/Triton and inference-serving frameworks.
vLLM, SGLang, TensorRT-LLM, Python, PyTorch, CUDA, Triton
1w
Save
Mark Applied
Hide
Senior Inference Engineer, AGI
Sunnyvale or Boston or Seattle or Los Angeles County
$193k-$262k/yr OnsiteFull Time
Amazon
AmazonNASDAQ: AMZN: Global online retail and cloud computing technology provider.
5+ YOERequires 5+ years software development, 4+ years systems architecture, a computer science bachelor's degree, 2+ years neural inference optimization, GPU optimization, real-time systems, and technical leadership experience.
Nsight Compute, Nsight Systems, vLLM, PyTorch, TensorRT-LLM, CUTLASS, Triton, CUDA, PTX, FlashAttention, NCCL, NVLink, AWS Neuron, Trainium, Microsoft Excel
1mo
Save
Mark Applied
Hide
AI Inference Engineer - Speech
Seattle or San Jose
$152k-$332k/yr HybridFull Time
Zoom
ZoomNasdaq: ZM: Provides a cloud-based platform for video, voice, and collaboration.
3+ YOEMaster's in CS/EE or related,3+ years in speech recognition or model inference,deep learning expertise,experience with Python,C/C++,CUDA,TensorRT,PyTorch,TensorFlow and GPU optimization.
Python, shell, C/C++, PyTorch, TensorFlow, CUDA, TensorRT, CUDA Graphs, NVIDIA GPUs, TPU, BrightHire
1mo
Save
Mark Applied
Hide
Research Engineer - LLM/VLM Inference Optimization (Seed Infra)
Seattle, Washington, United States
OnsiteFull Time
ByteDance
ByteDance: Developing AI-driven content platforms and mobile applications.
Bachelor's in CS/EE/Software, strong C/C++ and Python, experience with PyTorch or TensorFlow, production LLM/VLM inference optimization, GPU familiarity and operator optimization, containerization experience.
C, C++, Python, PyTorch, TensorFlow, CUDA, OpenCL, TensorRT, Triton, CUTLASS, FlashAttention, GEMM, GEMV, Conv2D
2mo
Save
Mark Applied
Hide
Fellow, AI Workload Optimization
Bellevue, Washington, United States
$224k-$384k/yr OnsiteFull Time
AMD
AMDNASDAQ: AMD: Designs and manufactures computer processors and graphics technology.
15+ YOE15+ years software development with 5+ years technical leadership; deep expertise in AI frameworks and ROCm; mastery of performance profiling and distributed training/inference optimization; PhD/Master's or equivalent experience.
PyTorch, JAX, vLLM, SGLang, ROCm, TorchProfiler, ROCm Profiler, Nsight
1w
Save
Mark Applied
Hide
Sr. Machine Learning Engineer, Foundation Models Inference - Cloud OS & Inference
Seattle, Washington, United States
OnsiteFull Time
Apple
AppleNASDAQ: AAPL: Designs and sells consumer electronics, software, and online services.
Build and optimize inference frameworks, services, and tools for large-scale foundation models, supporting low-latency AI services across Apple products.
2mo
Save
Mark Applied
Hide
Staff+ Software Engineer, Inference Runtime
San Francisco or Seattle or New York City
$405k-$485k/yr HybridFull Time
Anthropic
Anthropic: Developing safe and reliable artificial intelligence systems.
Senior IC with deep systems or ML infrastructure experience, hands-on performance profiling and optimization, accelerator ecosystem expertise (CUDA/TPU/Trainium), strong software engineering and cross-org alignment skills, and a relevant bachelor’s degree or equivalent.
Rust, Python, CUDA, XLA, Triton, NeuronX, AWS Neuron, Kubernetes, CI/CD
3mo
Save
Mark Applied
Hide
Staff Software Engineer, Inference
Sunnyvale or Bellevue
$188k-$275k/yr HybridFull Time
CoreWeave
CoreWeaveNASDAQ: CRWV: Cloud platform providing GPU-accelerated infrastructure for AI workloads.
8+ YOE8–12+ years in distributed systems; leadership of cross-team initiatives; proficient in Go, Python or C++; production Kubernetes expertise; low-latency, high-throughput system design and optimization.
Go, Python, C++, Kubernetes, CUDA, NCCL, RDMA, NUMA, TensorRT-LLM, Ray Serve, TorchServe
1mo
Save
Mark Applied
Hide
Machine Learning Engineer
Seattle, Washington, United States
$120k-$180k/yr OnsiteFull Time
Constellation Space
Constellation Space: AI-powered operating system for satellite network management.
BS/MS in CS or Engineering (or equivalent), proven experience deploying ML models to production, strong software engineering in Python and C++, experience with Docker, cloud platforms and MLOps, and optimizing low-latency inference.
Python, C++, Docker
3w
Save
Mark Applied
Hide
Staff Software Engineer - Data Cloud Applied ML
San Francisco or Seattle or New York City
$189k-$315k/yr OnsiteFull Time
Rippling
Rippling: Unified platform managing workforce HR, IT, and finance operations
8+ YOE8+ years software engineering experience, distributed systems ownership, experience training/deploying LLMs, model inference optimization, backend skills in Python/Go/Java, and cloud-native infrastructure (Kubernetes).
Python, Go, Java, Kubernetes
3w
Save
Mark Applied
Hide
Senior Machine Learning Engineer
Chicago or New York City or San Francisco or Seattle or Sunnyvale
$182k-$202k/yr OnsiteFull Time
Uber
UberNYSE: UBER: A technology platform for transportation, delivery, and freight.
4+ YOE4+ years building ML models; BS in CS/CE or related; experience with PyTorch, causal inference or constrained optimization preferred; product and marketplace experience a plus.
PyTorch
1mo
Save
Mark Applied
Hide
Principal Scientist - Data Pipeline Engineer
San Jose or Seattle or San Francisco
$206k-$388k/yr OnsiteFull Time
Adobe
AdobeNASDAQ: ADBE: Provides software for digital media creation and marketing analytics
10+ YOE10+ years in data engineering/ML infrastructure, distributed systems expertise, Python and a systems language, experience with Ray or Spark, GPU inference optimization, large-scale databases and data curation for model training.
Ray, Spark, Python, C++, Rust, Go, Java
2mo
Save
Mark Applied
Hide
Principal Software Engineering - AI Frameworks
Redmond or Mountain View or United States
$143k-$331k/yr HybridFull Time
Microsoft
MicrosoftNASDAQ: MSFT: Develops software, services, devices, and cloud computing solutions.
6+ YOEBachelor's in CS or related + 6+ years engineering experience (or equivalent); coding experience in C, C++, C#, Java, JavaScript, or Python; experience with inference stacks, performance optimization, open-source code; ability to pass Microsoft security screening.
C, C++, C#, Java, JavaScript, Python, ONNX, ONNX Runtime, Foundry Local, VSCode, SQL Server, CLI, SDK, REST API