28 inference optimization engineer jobs at 16 companies in Federal Way, WA

1mo
Save
Mark Applied
Hide
Senior Inference Engineer, GPU Kernel Optimization
Santa Clara or Austin or New York City or Seattle
$184k-$288k/yr OnsiteFull Time
NVIDIA
NVIDIANASDAQ: NVDA: Computing platform for AI and accelerated graphics.
6+ YOEMaster's or PhD in computer science, computer engineering, or related field, 6+ years' experience, strong Python and C++, GPU profiling, LLM inference frameworks, and CUDA kernel optimization expertise.
Python, C++, CUPTI, NSYS, NCU, TRT-LLM, SGLang, vLLM, CUDA, CUTLASS, Triton, PTX, SASS, LLVM, MLIR, ptxas, FlashInfer
1mo
Save
Mark Applied
Hide
Senior Inference Engineer, GPU Kernel Optimization
Santa Clara or Austin or New York City or Seattle
$184k-$288k/yr OnsiteFull Time
NVIDIA
NVIDIANASDAQ: NVDA: Computing platform for AI and accelerated graphics.
6+ YOE6+ years industry experience; strong Python and C++; hands-on GPU profiling (CUPTI, NSYS, NCU); experience with LLM inference frameworks and GPU kernel optimization; advanced degree or equivalent experience.
Python, C++, CUPTI, NSYS, NCU, TRT-LLM, SGLang, vLLM, CUDA, CUTLASS, Triton, PTX, SASS, LLVM, MLIR, ptxas, FlashInfer
1mo
Save
Mark Applied
Hide
Senior Inference Engineer, GPU Kernel Optimization
Santa Clara or Austin or New York City or Seattle
$184k-$288k/yr OnsiteFull Time
NVIDIA
NVIDIANASDAQ: NVDA: Computing platform for AI and accelerated graphics.
6+ YOEMaster's/PhD or equivalent,6+ years industry experience,agentic AI systems experience,strong Python/C++,GPU profiling (CUPTI,NSYS,NCU),LLM inference frameworks,CUDA/CUTLASS/Triton and PTX/SASS familiarity.
Python, C++, CUPTI, NSYS, NCU, TRT-LLM, SGLang, vLLM, CUDA, CUTLASS, Triton, PTX, SASS, LLVM, MLIR, ptxas
2mo
Save
Mark Applied
Hide
Staff Engineer, Inference Optimizations
Seattle, Washington, United States
$191k-$239k/yr HybridFull Time
DigitalOcean
DigitalOceanNYSE: DOCN: The AI-Native Cloud purpose-built for inference and agentic workloads.
8+ YOE8+ years building/operating multi-tenant distributed systems, 2+ years Go/Golang, 2+ years Kubernetes, strong SRE/observability and production debugging skills, experience optimizing inference and GPU utilization.
Go, Golang, Kubernetes, vLLM, Triton, TensorRT-LLM
3d
Save
Mark Applied
Hide
Senior Inference Engineer, AGI
Sunnyvale or Boston or Seattle or Los Angeles County
$167k-$260k/yr OnsiteFull Time
Amazon
AmazonNASDAQ: AMZN: Multinational technology focused on e-commerce and cloud computing.
3+ YOERequires 3+ years building machine learning models, 2+ years optimizing neural-model inference, production real-time inference experience, GPU optimization expertise, and Java, C++, or Python programming.
Java, C++, Python, R, scikit-learn, Spark MLLib, MxNet, Tensorflow, numpy, scipy, vLLM, PyTorch, TensorRT-LLM, CUTLASS, Triton, CUDA, PTX, FlashAttention, NCCL, NVLink, AWS Neuron, Trainium, NVIDIA GPU
2mo
Save
Mark Applied
Hide
Member of Technical Staff — Model Optimization and Inference
Seattle, Washington, United States
$250k-$350k/yr OnsiteFull Time
Nuance Labs
Nuance Labs: Private AI research building real-time audiovisual foundation models for face-to-face conversational AI.
Deep expertise in LLM and diffusion-model inference optimization, KV cache strategies, quantization (INT8/INT4, GPTQ/AWQ), profiling/benchmarking, and strong Python/PyTorch skills; familiarity with CUDA/Triton and inference-serving frameworks.
vLLM, SGLang, TensorRT-LLM, Python, PyTorch, CUDA, Triton
3d
Save
Mark Applied
Hide
AI Inference Engineer
San Jose or Seattle or United States
$177k-$265k/yr OnsiteFull Time
F5
F5NASDAQ: FFIV: Delivering and securing applications across any multi-cloud environment.
Proficiency in Python, C++, Rust, or Golang; experience with inference tools and cloud infrastructure; expertise in GPU and AI hardware optimization; scalable AI serving and monitoring experience.
vLLM, TGI (Text Generation Inference), NVIDIA Triton, NVIDIA GPUs, CUDA, TensorRT, Apple Silicon, CoreML, TPUs, LPUs, Kubernetes, Python, C++, Rust, Golang, Llama.cpp, Ollama, Docker, AWS, GCP, Azure, Speculative Decoding, PagedAttention, Triton kernels, MLOps, SRE, Time to First Token (TTFT), SLAs
1mo
Save
Mark Applied
Hide
AI Inference Engineer - Speech
Seattle or San Jose
$152k-$332k/yr HybridFull Time
Zoom
ZoomNasdaq Global Select Market: ZM: American publicly traded communications platform serving businesses and individuals with AI-assisted video, voice, chat, and phone services.
3+ YOEMaster's in CS/EE or related,3+ years in speech recognition or model inference,deep learning expertise,experience with Python,C/C++,CUDA,TensorRT,PyTorch,TensorFlow and GPU optimization.
Python, shell, C/C++, PyTorch, TensorFlow, CUDA, TensorRT, CUDA Graphs, NVIDIA GPUs, TPU, BrightHire
1mo
Save
Mark Applied
Hide
Research Engineer - LLM/VLM Inference Optimization (Seed Infra)
Seattle, Washington, United States
OnsiteFull Time
ByteDance
ByteDance: Global technology specializing in AI-powered content platforms.
Bachelor's in CS/EE/Software, strong C/C++ and Python, experience with PyTorch or TensorFlow, production LLM/VLM inference optimization, GPU familiarity and operator optimization, containerization experience.
C, C++, Python, PyTorch, TensorFlow, CUDA, OpenCL, TensorRT, Triton, CUTLASS, FlashAttention, GEMM, GEMV, Conv2D
2mo
Save
Mark Applied
Hide
Fellow, AI Workload Optimization
Bellevue, Washington, United States
$224k-$384k/yr OnsiteFull Time
AMD
AMDNASDAQ: AMD: Leader in high-performance computing, graphics, and visualization technologies.
15+ YOE15+ years software development with 5+ years technical leadership; deep expertise in AI frameworks and ROCm; mastery of performance profiling and distributed training/inference optimization; PhD/Master's or equivalent experience.
PyTorch, JAX, vLLM, SGLang, ROCm, TorchProfiler, ROCm Profiler, Nsight
2w
Save
Mark Applied
Hide
Sr. Machine Learning Engineer, Foundation Models Inference - Cloud OS & Inference
Seattle, Washington, United States
OnsiteFull Time
Apple
AppleNASDAQ: AAPL: Designing and manufacturing consumer electronics, software, and digital services.
Build and optimize inference frameworks, services, and tools for large-scale foundation models, supporting low-latency AI services across Apple products.
2mo
Save
Mark Applied
Hide
Staff+ Software Engineer, Inference Runtime
San Francisco or Seattle or New York City
$405k-$485k/yr HybridFull Time
Anthropic
Anthropic: AI research developing safe and steerable AI systems.
Senior IC with deep systems or ML infrastructure experience, hands-on performance profiling and optimization, accelerator ecosystem expertise (CUDA/TPU/Trainium), strong software engineering and cross-org alignment skills, and a relevant bachelor’s degree or equivalent.
Rust, Python, CUDA, XLA, Triton, NeuronX, AWS Neuron, Kubernetes, CI/CD
3mo
Save
Mark Applied
Hide
Staff Software Engineer, Inference
Sunnyvale or Bellevue
$188k-$275k/yr HybridFull Time
CoreWeave
CoreWeaveNasdaq: CRWV: Specialized cloud provider for large-scale AI and machine learning.
8+ YOE8–12+ years in distributed systems; leadership of cross-team initiatives; proficient in Go, Python or C++; production Kubernetes expertise; low-latency, high-throughput system design and optimization.
Go, Python, C++, Kubernetes, CUDA, NCCL, RDMA, NUMA, TensorRT-LLM, Ray Serve, TorchServe
2mo
Save
Mark Applied
Hide
Machine Learning Engineer
Seattle, Washington, United States
$120k-$180k/yr OnsiteFull Time
Constellation Space
Constellation Space: Private Seattle software building AI-driven satellite-network operations infrastructure for space companies.
BS/MS in CS or Engineering (or equivalent), proven experience deploying ML models to production, strong software engineering in Python and C++, experience with Docker, cloud platforms and MLOps, and optimizing low-latency inference.
Python, C++, Docker
4w
Save
Mark Applied
Hide
Staff Software Engineer - Data Cloud Applied ML
San Francisco or Seattle or New York City
$189k-$315k/yr OnsiteFull Time
Rippling
Rippling: Unified workforce management platform for HR, IT, and Finance.
8+ YOE8+ years software engineering experience, distributed systems ownership, experience training/deploying LLMs, model inference optimization, backend skills in Python/Go/Java, and cloud-native infrastructure (Kubernetes).
Python, Go, Java, Kubernetes
1mo
Save
Mark Applied
Hide
Senior Machine Learning Engineer
Chicago or New York City or San Francisco or Seattle or Sunnyvale
$182k-$202k/yr OnsiteFull Time
Uber Freight
Uber FreightNYSE: UBER: Global technology platform for ride-hailing, delivery, and freight logistics.
4+ YOE4+ years building ML models; BS in CS/CE or related; experience with PyTorch, causal inference or constrained optimization preferred; product and marketplace experience a plus.
PyTorch
1mo
Save
Mark Applied
Hide
Principal Scientist - Data Pipeline Engineer
San Jose or Seattle or San Francisco
$206k-$388k/yr OnsiteFull Time
Adobe
AdobeNASDAQ: ADBE: Empowering everyone to create through innovative digital experiences.
10+ YOE10+ years in data engineering/ML infrastructure, distributed systems expertise, Python and a systems language, experience with Ray or Spark, GPU inference optimization, large-scale databases and data curation for model training.
Ray, Spark, Python, C++, Rust, Go, Java
2mo
Save
Mark Applied
Hide
Principal Software Engineering - AI Frameworks
Redmond or Mountain View or United States
$143k-$331k/yr HybridFull Time
Microsoft
MicrosoftNASDAQ: MSFT: Multinational technology providing software, cloud, and AI solutions.
6+ YOEBachelor's in CS or related + 6+ years engineering experience (or equivalent); coding experience in C, C++, C#, Java, JavaScript, or Python; experience with inference stacks, performance optimization, open-source code; ability to pass Microsoft security screening.
C, C++, C#, Java, JavaScript, Python, ONNX, ONNX Runtime, Foundry Local, VSCode, SQL Server, CLI, SDK, REST API