9 model serving engineer jobs at 7 companies in Washington
2mo
Save
Mark Applied
Hide
2mo
Member of Technical Staff — Model Optimization and Inference
Seattle, Washington, United States
$250k-$350k/yrOnsiteFull Time
Nuance Labs: Private AI research building real-time audiovisual foundation models for face-to-face conversational AI.
Deep expertise in LLM and diffusion-model inference optimization, KV cache strategies, quantization (INT8/INT4, GPTQ/AWQ), profiling/benchmarking, and strong Python/PyTorch skills; familiarity with CUDA/Triton and inference-serving frameworks.
Distributed Systems Engineer 5 - Core Ad Serving Platform
New York City or Seattle or Los Angeles or Los Gatos
$388k-$619k/yrOnsiteFull Time
NetflixNASDAQ: NFLX: Global subscription-based streaming entertainment service and content producer.
7+ YOE7+ years experience with at least 4+ years in Ads domain, expertise building and operating large-scale distributed systems, ad-server components, API and data model design, SLO-driven development, and incident response.
AtlassianNASDAQ: TEAM: Unleashing the potential of every team.
Design, develop, and deploy production ML systems for search and retrieval; ensure low-latency, high-availability serving; integrate models with Triton and PyTorch; drive observability, cost optimization, and mentor engineers.
New York City or Seattle or Los Angeles or Los Gatos
$466k-$750k/yrOnsiteFull Time
NetflixNASDAQ: NFLX: Global subscription-based streaming entertainment service and content producer.
7+ YOE7+ years software engineering experience with 3+ years on ML infrastructure or model serving; proficiency in Java, Python, or Scala; experience building high‑QPS, low‑latency model serving, feature serving, and model monitoring.
NetflixNASDAQ: NFLX: Global subscription-based streaming entertainment service and content producer.
7+ YOE7+ years software engineering; 3+ years ML infrastructure, model serving, or ML platform experience in ads/real-time decisioning; real-time model serving with sub-20ms latency; proficiency in Java, Python, or Scala; experience with ML serving frameworks and real-time feature pipelines; strong model monitoring and production readiness.
Java, Python, Scala, ML serving frameworks, feature stores, model registries
UnityNYSE: U: Develops tools for games and interactive experiences.
Experience building and operating ML infrastructure and model serving systems; proficiency in Golang or Python; Kubernetes experience; familiarity with ML serving frameworks and distributed systems; strong collaboration and communication.
Livingston or New York or Sunnyvale or San Francisco or Bellevue
$207k-$275k/yrOnsiteFull Time
CoreWeaveNasdaq: CRWV: Specialized cloud provider for large-scale AI and machine learning.
10+ YOE10+ years in distributed systems/ML infrastructure or production AI engineering; 5+ years with AI runtime systems; expertise in model serving, batching, execution isolation, GPU memory management, and Kubernetes; strong customer-facing and commercial skills.