8 model serving engineer jobs at 6 companies in Lake Stevens, WA
2mo
Save
Mark Applied
Hide
2mo
Member of Technical Staff — Model Optimization and Inference
Seattle, Washington, United States
$250k-$350k/yrOnsiteFull Time
Nuance Labs: A building photorealistic, real-time AI avatars and full-duplex audiovisual systems.
Deep expertise in LLM and diffusion-model inference optimization, KV cache strategies, quantization (INT8/INT4, GPTQ/AWQ), profiling/benchmarking, and strong Python/PyTorch skills; familiarity with CUDA/Triton and inference-serving frameworks.
Distributed Systems Engineer 5 - Core Ad Serving Platform
New York City or Seattle or Los Angeles or Los Gatos
$388k-$619k/yrOnsiteFull Time
NetflixNASDAQ: NFLX: Provider of global streaming entertainment and video content.
7+ YOE7+ years experience with at least 4+ years in Ads domain, expertise building and operating large-scale distributed systems, ad-server components, API and data model design, SLO-driven development, and incident response.
AtlassianNASDAQ: TEAM: Develops software for team collaboration and project management.
Design, develop, and deploy production ML systems for search and retrieval; ensure low-latency, high-availability serving; integrate models with Triton and PyTorch; drive observability, cost optimization, and mentor engineers.
New York City or Seattle or Los Angeles or Los Gatos
$466k-$750k/yrOnsiteFull Time
NetflixNASDAQ: NFLX: Provider of global streaming entertainment and video content.
7+ YOE7+ years software engineering experience with 3+ years on ML infrastructure or model serving; proficiency in Java, Python, or Scala; experience building high‑QPS, low‑latency model serving, feature serving, and model monitoring.
NetflixNASDAQ: NFLX: Global video streaming and media production service.
7+ YOE7+ years software engineering; 3+ years ML infrastructure, model serving, or ML platform experience in ads/real-time decisioning; real-time model serving with sub-20ms latency; proficiency in Java, Python, or Scala; experience with ML serving frameworks and real-time feature pipelines; strong model monitoring and production readiness.
Java, Python, Scala, ML serving frameworks, feature stores, model registries
Arlington or Atlanta or Boston or Chicago or Culver City or Houston or Irving or Morristown or New York City or San Diego or San Francisco or Seattle
$59k-$188k/yrHybridFull Time
AccentureNYSE: ACN: Global professional services firm providing consulting and technology solutions.
3+ YOEBachelor's degree in a relevant field; 3+ years of client-facing technical delivery and hands-on backend, API, full-stack, data, model-serving, or agentic orchestration experience; Python proficiency required.
Python, TypeScript, JavaScript, Electron, FastAPI, Docker, Microsoft Windows, MCP, Microsoft Azure
Livingston or New York or Sunnyvale or San Francisco or Bellevue
$207k-$275k/yrOnsiteFull Time
CoreWeaveNASDAQ: CRWV: Cloud platform providing GPU-accelerated infrastructure for AI workloads.
10+ YOE10+ years in distributed systems/ML infrastructure or production AI engineering; 5+ years with AI runtime systems; expertise in model serving, batching, execution isolation, GPU memory management, and Kubernetes; strong customer-facing and commercial skills.