8 model serving engineer jobs at 6 companies in Lake Stevens, WA

2mo
Save
Mark Applied
Hide
Member of Technical Staff — Model Optimization and Inference
Seattle, Washington, United States
$250k-$350k/yr OnsiteFull Time
Nuance Labs
Nuance Labs: A building photorealistic, real-time AI avatars and full-duplex audiovisual systems.
Deep expertise in LLM and diffusion-model inference optimization, KV cache strategies, quantization (INT8/INT4, GPTQ/AWQ), profiling/benchmarking, and strong Python/PyTorch skills; familiarity with CUDA/Triton and inference-serving frameworks.
vLLM, SGLang, TensorRT-LLM, Python, PyTorch, CUDA, Triton
1w
Save
Mark Applied
Hide
Machine Learning Engineer, Foundation Model Services
Seattle, Washington, United States
OnsiteFull Time
Apple
AppleNASDAQ: AAPL: Designs and sells consumer electronics, software, and online services.
Build frameworks, services, and tools for production foundation models, optimizing and serving large language, vision, and speech models at scale.
3w
Save
Mark Applied
Hide
Distributed Systems Engineer 5 - Core Ad Serving Platform
New York City or Seattle or Los Angeles or Los Gatos
$388k-$619k/yr OnsiteFull Time
Netflix
NetflixNASDAQ: NFLX: Provider of global streaming entertainment and video content.
7+ YOE7+ years experience with at least 4+ years in Ads domain, expertise building and operating large-scale distributed systems, ad-server components, API and data model design, SLO-driven development, and incident response.
3w
Save
Mark Applied
Hide
Senior Machine Learning System Engineer
Seattle, Washington, United States
$149k-$235k/yr HybridFull Time
Atlassian
AtlassianNASDAQ: TEAM: Develops software for team collaboration and project management.
Design, develop, and deploy production ML systems for search and retrieval; ensure low-latency, high-availability serving; integrate models with Triton and PyTorch; drive observability, cost optimization, and mentor engineers.
Triton, PyTorch
2mo
Save
Mark Applied
Hide
Machine Learning Engineer 5 - Decisioning & Optimization
New York City or Seattle or Los Angeles or Los Gatos
$466k-$750k/yr OnsiteFull Time
Netflix
NetflixNASDAQ: NFLX: Provider of global streaming entertainment and video content.
7+ YOE7+ years software engineering experience with 3+ years on ML infrastructure or model serving; proficiency in Java, Python, or Scala; experience building high‑QPS, low‑latency model serving, feature serving, and model monitoring.
Java, Python, Scala, Chronon, Signal Service, JVM
2mo
Save
Mark Applied
Hide
Machine Learning Engineer 5 - Decisioning & Optimization
New York or Los Angeles or Los Gatos or Seattle
$466k-$750k/yr OnsiteFull Time
Netflix
NetflixNASDAQ: NFLX: Global video streaming and media production service.
7+ YOE7+ years software engineering; 3+ years ML infrastructure, model serving, or ML platform experience in ads/real-time decisioning; real-time model serving with sub-20ms latency; proficiency in Java, Python, or Scala; experience with ML serving frameworks and real-time feature pipelines; strong model monitoring and production readiness.
Java, Python, Scala, ML serving frameworks, feature stores, model registries
4d
Save
Mark Applied
Hide
Trading, Investment & Optimization - QuantAI Engineer (Hybrid)
Arlington or Atlanta or Boston or Chicago or Culver City or Houston or Irving or Morristown or New York City or San Diego or San Francisco or Seattle
$59k-$188k/yr HybridFull Time
Accenture
AccentureNYSE: ACN: Global professional services firm providing consulting and technology solutions.
3+ YOEBachelor's degree in a relevant field; 3+ years of client-facing technical delivery and hands-on backend, API, full-stack, data, model-serving, or agentic orchestration experience; Python proficiency required.
Python, TypeScript, JavaScript, Electron, FastAPI, Docker, Microsoft Windows, MCP, Microsoft Azure
1mo
Save
Mark Applied
Hide
Solution Specialist, AI Runtime Services
Livingston or New York or Sunnyvale or San Francisco or Bellevue
$207k-$275k/yr OnsiteFull Time
CoreWeave
CoreWeaveNASDAQ: CRWV: Cloud platform providing GPU-accelerated infrastructure for AI workloads.
10+ YOE10+ years in distributed systems/ML infrastructure or production AI engineering; 5+ years with AI runtime systems; expertise in model serving, batching, execution isolation, GPU memory management, and Kubernetes; strong customer-facing and commercial skills.
vLLM, TensorRT-LLM, TGI, Triton, Kubernetes, Inference, Sandboxes