85 inference engineer jobs at 31 companies in Lynnwood, WA

1mo
Save
Mark Applied
Hide
Staff Engineer, Inference Optimizations
Seattle, Washington, United States
$191k-$239k/yr HybridFull Time
DigitalOcean
DigitalOceanNew York Stock Exchange: DOCN: Simplifies cloud infrastructure for developers, startups, and SMBs.
8+ YOE8+ years building/operating multi-tenant distributed systems, 2+ years Go/Golang, 2+ years Kubernetes, strong SRE/observability and production debugging skills, experience optimizing inference and GPU utilization.
Go, Golang, Kubernetes, vLLM, Triton, TensorRT-LLM
3w
Save
Mark Applied
Hide
Senior Inference Engineer, GPU Kernel Optimization
Santa Clara or Austin or New York City or Seattle
$184k-$288k/yr OnsiteFull Time
NVIDIA
NVIDIANASDAQ: NVDA: Designs GPU-accelerated computing and artificial intelligence hardware.
6+ YOE6+ years industry experience; strong Python and C++; hands-on GPU profiling (CUPTI, NSYS, NCU); experience with LLM inference frameworks and GPU kernel optimization; advanced degree or equivalent experience.
Python, C++, CUPTI, NSYS, NCU, TRT-LLM, SGLang, vLLM, CUDA, CUTLASS, Triton, PTX, SASS, LLVM, MLIR, ptxas, FlashInfer
3w
Save
Mark Applied
Hide
Senior Inference Engineer, GPU Kernel Optimization
Santa Clara or Austin or New York City or Seattle
$184k-$288k/yr OnsiteFull Time
NVIDIA
NVIDIANASDAQ: NVDA: Designs graphics processing units and artificial intelligence hardware.
6+ YOEMaster's/PhD or equivalent,6+ years industry experience,agentic AI systems experience,strong Python/C++,GPU profiling (CUPTI,NSYS,NCU),LLM inference frameworks,CUDA/CUTLASS/Triton and PTX/SASS familiarity.
Python, C++, CUPTI, NSYS, NCU, TRT-LLM, SGLang, vLLM, CUDA, CUTLASS, Triton, PTX, SASS, LLVM, MLIR, ptxas
1d
Save
Mark Applied
Hide
Senior Inference Engineer, AGI
Sunnyvale or Boston or Seattle or Los Angeles County
$193k-$262k/yr OnsiteFull Time
Amazon
AmazonNASDAQ: AMZN: Global online retail and cloud computing technology provider.
5+ YOERequires 5+ years software development, 4+ years systems architecture, a computer science bachelor's degree, 2+ years neural inference optimization, GPU optimization, real-time systems, and technical leadership experience.
Nsight Compute, Nsight Systems, vLLM, PyTorch, TensorRT-LLM, CUTLASS, Triton, CUDA, PTX, FlashAttention, NCCL, NVLink, AWS Neuron, Trainium, Microsoft Excel
1mo
Save
Mark Applied
Hide
AI Inference Engineer - Speech
Seattle or San Jose
$152k-$332k/yr HybridFull Time
Zoom
ZoomNasdaq: ZM: Provides a cloud-based platform for video, voice, and collaboration.
3+ YOEMaster's in CS/EE or related,3+ years in speech recognition or model inference,deep learning expertise,experience with Python,C/C++,CUDA,TensorRT,PyTorch,TensorFlow and GPU optimization.
Python, shell, C/C++, PyTorch, TensorFlow, CUDA, TensorRT, CUDA Graphs, NVIDIA GPUs, TPU, BrightHire
1w
Save
Mark Applied
Hide
AI Inference Platform Engineer
Seattle, Washington, United States
OnsiteFull Time
Apple
AppleNASDAQ: AAPL: Designs and sells consumer electronics, software, and online services.
Build tooling, automation, benchmarking systems, capacity models, and data analysis pipelines for AI inference infrastructure and scaling decisions.
3mo
Save
Mark Applied
Hide
AI Inference Infrastructure Software Engineer (Kubernetes / Cloud)
Seattle, Washington, United States
HybridFull Time
ElastixAI
ElastixAI: A software startup focused on AI inference infrastructure and high-performance cloud platforms.
3+ YOE3–5 years in Kubernetes and cloud deployments; BS in CS/Software Eng; strong Python/Bash/Go; experience with Terraform/Pulumi, Ansible/Puppet, GitOps; production ML/inference on Kubernetes.
Kubernetes, AWS, GCP, Terraform, Pulumi, Ansible, GitOps, Argo CD, Flux, Python, Bash, Go, Linux
3mo
Save
Mark Applied
Hide
Performance Engineer, Inference Systems
San Francisco or New York City or Seattle
$350k-$850k/yr OnsiteFull Time
Anthropic
Anthropic: Developing safe and reliable artificial intelligence systems.
Hands-on performance engineering with Python, data analysis, and cross-layer investigations; strong communication of quantitative results.
Python, SQL, Pandas
10h
Save
Mark Applied
Hide
Staff Machine Learning Engineer, Causal Inference
San Francisco or Sunnyvale or Los Angeles or Seattle or New York City
$204k-$299k/yr OnsiteFull Time
DoorDash
DoorDashNASDAQ: DASH: On-demand delivery platform connecting consumers with local merchants.
Deep causal inference, econometrics, experimentation, or causal ML experience; production ML engineering; rigorous model evaluation; and ability to collaborate across engineering, analytics, product, and business teams.
ML, CUPED, Covey, Covey Scout for Inbound
2w
Save
Mark Applied
Hide
Staff Machine Learning Engineer, Causal Inference
San Francisco or Sunnyvale or Los Angeles or Seattle or New York City
$204k-$299k/yr OnsiteFull Time
DoorDash
DoorDashNYSE: DASH: Local food delivery and on-demand logistics platform.
Requires practical causal inference, econometrics, experimentation, production ML engineering, reliable pipeline and model development, rigorous evaluation, and cross-functional product judgment.
Machine Learning (ML), CUPED, contextual bandits, off-policy evaluation, Covey Scout for Inbound
2mo
Save
Mark Applied
Hide
Member of Technical Staff — Model Optimization and Inference
Seattle, Washington, United States
$250k-$350k/yr OnsiteFull Time
Nuance Labs
Nuance Labs: A building photorealistic, real-time AI avatars and full-duplex audiovisual systems.
Deep expertise in LLM and diffusion-model inference optimization, KV cache strategies, quantization (INT8/INT4, GPTQ/AWQ), profiling/benchmarking, and strong Python/PyTorch skills; familiarity with CUDA/Triton and inference-serving frameworks.
vLLM, SGLang, TensorRT-LLM, Python, PyTorch, CUDA, Triton
1mo
Save
Mark Applied
Hide
Software Engineer Graduate (Inference Infrastructure) - 2026 Start (PhD)
Seattle, Washington, United States
OnsiteFull Time
ByteDance
ByteDance: Developing AI-driven content platforms and mobile applications.
PhD in CS or related discipline, strong knowledge of large-model inference, distributed systems, container orchestration, and proficiency in Go/Rust/Python/C++ with cloud/ML infrastructure experience.
AIBrix, Kubernetes, vLLM, SGLang, TensorRT-LLM, Docker, Go, Rust, Python, C++, Ray, CUDA, AWS, Azure, GCP, SageMaker, Azure ML, Vertex AI, DeepSpeed, PyTorch
3mo
Save
Mark Applied
Hide
Staff Software Engineer, Inference
Sunnyvale or Bellevue
$188k-$275k/yr HybridFull Time
CoreWeave
CoreWeaveNASDAQ: CRWV: Cloud platform providing GPU-accelerated infrastructure for AI workloads.
8+ YOE8–12+ years in distributed systems; leadership of cross-team initiatives; proficient in Go, Python or C++; production Kubernetes expertise; low-latency, high-throughput system design and optimization.
Go, Python, C++, Kubernetes, CUDA, NCCL, RDMA, NUMA, TensorRT-LLM, Ray Serve, TorchServe
1mo
Save
Mark Applied
Hide
Machine Learning Engineer
Seattle, Washington, United States
$120k-$180k/yr OnsiteFull Time
Constellation Space
Constellation Space: AI-powered operating system for satellite network management.
BS/MS in CS or Engineering (or equivalent), proven experience deploying ML models to production, strong software engineering in Python and C++, experience with Docker, cloud platforms and MLOps, and optimizing low-latency inference.
Python, C++, Docker
1mo
Save
Mark Applied
Hide
Senior Staff Engineer, GDC AI Inference Platform
Sunnyvale or Kirkland
$262k-$365k/yr OnsiteFull Time
Google
GoogleNASDAQ: GOOGL: Provides online search, advertising, cloud computing, and consumer electronics.
8+ YOEBachelor's or equivalent; 8+ years software development with C++, Java, Python, Kotlin, or Go; 4+ years technical leadership; 5+ years testing/launch experience; EMR not applicable; strong communication and critical thinking.
C++, Java, Python, Kotlin, Go, Kubernetes, Docker
1mo
Save
Mark Applied
Hide
Machine Learning Platform Engineer
San Francisco or New York City or Los Angeles or Seattle
$245k-$345k/yr HybridFull Time
Whatnot
Whatnot: Social marketplace for buying and selling via live streams
4+ YOE4+ years building ML systems, 3+ years engineering production systems, 1+ year Python, experience with distributed training/inference, databases, monitoring, and cloud services.
Python, PostgreSQL, DynamoDB, Elasticsearch, Redis, DataDog, Grafana, AWS Sagemaker, Lambda, Kinesis, S3, EC2, EKS, ECS, Apache Kafka, Flink
2w
Save
Mark Applied
Hide
Staff Software Engineer - Data Cloud Applied ML
San Francisco or Seattle or New York City
$189k-$315k/yr OnsiteFull Time
Rippling
Rippling: Unified platform managing workforce HR, IT, and finance operations
8+ YOE8+ years software engineering experience, distributed systems ownership, experience training/deploying LLMs, model inference optimization, backend skills in Python/Go/Java, and cloud-native infrastructure (Kubernetes).
Python, Go, Java, Kubernetes
3mo
Save
Mark Applied
Hide
Senior Solutions Engineer
Bellevue or Palo Alto
OnsiteFull Time
Harell Data
Harell Data: A managed platform that enables secure sharing of scientific datasets and provides compute and tooling for model training and deployment.
5+ YOE5+ years in customer-facing technical roles; familiar with ML training/inference, cloud infrastructure (AWS/GCP), strong written/verbal communication, ability to create docs and runbooks.
PyTorch, Hugging Face, AWS, GCP
2w
Save
Mark Applied
Hide
Senior Machine Learning Engineer
Chicago or New York City or San Francisco or Seattle or Sunnyvale
$182k-$202k/yr OnsiteFull Time
Uber
UberNYSE: UBER: A technology platform for transportation, delivery, and freight.
4+ YOE4+ years building ML models; BS in CS/CE or related; experience with PyTorch, causal inference or constrained optimization preferred; product and marketplace experience a plus.
PyTorch
1mo
Save
Mark Applied
Hide
Member of Technical Staff - Machine Learning Infrastructure Engineer
San Francisco or Toronto or Seattle
$180k-$300k/yr OnsiteFull Time
Preference Model
Preference Model: Building reinforcement learning environments to train frontier AI models.
Experienced software engineer with production ML/data infrastructure skills, proficiency with PyTorch or JAX, distributed systems, AWS/GCP, Kubernetes, data pipelines, and familiarity with transformers and inference libraries like vLLM.
PyTorch, JAX, AWS, GCP, Kubernetes, transformers, vLLM, SGLang