9 model serving engineer jobs at 7 companies in Washington

2mo
Save
Mark Applied
Hide
Member of Technical Staff — Model Optimization and Inference
Seattle, Washington, United States
$250k-$350k/yr OnsiteFull Time
Nuance Labs
Nuance Labs: Private AI research building real-time audiovisual foundation models for face-to-face conversational AI.
Deep expertise in LLM and diffusion-model inference optimization, KV cache strategies, quantization (INT8/INT4, GPTQ/AWQ), profiling/benchmarking, and strong Python/PyTorch skills; familiarity with CUDA/Triton and inference-serving frameworks.
vLLM, SGLang, TensorRT-LLM, Python, PyTorch, CUDA, Triton
2w
Save
Mark Applied
Hide
Machine Learning Engineer, Foundation Model Services
Seattle, Washington, United States
OnsiteFull Time
Apple
AppleNASDAQ: AAPL: Designing and manufacturing consumer electronics, software, and digital services.
Build frameworks, services, and tools for production foundation models, optimizing and serving large language, vision, and speech models at scale.
1mo
Save
Mark Applied
Hide
Distributed Systems Engineer 5 - Core Ad Serving Platform
New York City or Seattle or Los Angeles or Los Gatos
$388k-$619k/yr OnsiteFull Time
Netflix
NetflixNASDAQ: NFLX: Global subscription-based streaming entertainment service and content producer.
7+ YOE7+ years experience with at least 4+ years in Ads domain, expertise building and operating large-scale distributed systems, ad-server components, API and data model design, SLO-driven development, and incident response.
1mo
Save
Mark Applied
Hide
Senior Machine Learning System Engineer
Seattle, Washington, United States
$149k-$235k/yr HybridFull Time
Atlassian
AtlassianNASDAQ: TEAM: Unleashing the potential of every team.
Design, develop, and deploy production ML systems for search and retrieval; ensure low-latency, high-availability serving; integrate models with Triton and PyTorch; drive observability, cost optimization, and mentor engineers.
Triton, PyTorch
3mo
Save
Mark Applied
Hide
Machine Learning Engineer 5 - Decisioning & Optimization
New York City or Seattle or Los Angeles or Los Gatos
$466k-$750k/yr OnsiteFull Time
Netflix
NetflixNASDAQ: NFLX: Global subscription-based streaming entertainment service and content producer.
7+ YOE7+ years software engineering experience with 3+ years on ML infrastructure or model serving; proficiency in Java, Python, or Scala; experience building high‑QPS, low‑latency model serving, feature serving, and model monitoring.
Java, Python, Scala, Chronon, Signal Service, JVM
3mo
Save
Mark Applied
Hide
Machine Learning Engineer 5 - Decisioning & Optimization
New York or Los Angeles or Los Gatos or Seattle
$466k-$750k/yr OnsiteFull Time
Netflix
NetflixNASDAQ: NFLX: Global subscription-based streaming entertainment service and content producer.
7+ YOE7+ years software engineering; 3+ years ML infrastructure, model serving, or ML platform experience in ads/real-time decisioning; real-time model serving with sub-20ms latency; proficiency in Java, Python, or Scala; experience with ML serving frameworks and real-time feature pipelines; strong model monitoring and production readiness.
Java, Python, Scala, ML serving frameworks, feature stores, model registries
2mo
Save
Mark Applied
Hide
Machine Learning Engineer, New Grad
Mountain View or Bellevue or San Francisco
$210k-$316k/yr OnsiteFull Time
Unity
UnityNYSE: U: Develops tools for games and interactive experiences.
Experience building and operating ML infrastructure and model serving systems; proficiency in Golang or Python; Kubernetes experience; familiarity with ML serving frameworks and distributed systems; strong collaboration and communication.
Golang, Python, Kubernetes, Ray Serve, Triton, TorchServe, Terraform, Prometheus, Grafana, OpenTelemetry
2mo
Save
Mark Applied
Hide
Sr Software Engineer, MLOps (VIRTUAL, WA, US, 00000)
Washington, United States
$150k-$180k/yr RemoteFull Time
Vivint Smart Home
Vivint Smart Home: U.S. smart home security and automation serving homeowners with professionally installed, monitored connected home systems.
2+ YOEBachelor's (CS/SE/AI/related) +5 yrs or Master's +2 yrs; experience building production ML platforms, model serving, Python, cloud engineering, CI/CD, Git, IaC, monitoring, DVC; strong cross-team communication.
Python, Git, CI/CD, infrastructure-as-code, DVC, GCP, AWS, Cloud Run, Kubernetes, Vertex AI, SageMaker, MLflow
2mo
Save
Mark Applied
Hide
Solution Specialist, AI Runtime Services
Livingston or New York or Sunnyvale or San Francisco or Bellevue
$207k-$275k/yr OnsiteFull Time
CoreWeave
CoreWeaveNasdaq: CRWV: Specialized cloud provider for large-scale AI and machine learning.
10+ YOE10+ years in distributed systems/ML infrastructure or production AI engineering; 5+ years with AI runtime systems; expertise in model serving, batching, execution isolation, GPU memory management, and Kubernetes; strong customer-facing and commercial skills.
vLLM, TensorRT-LLM, TGI, Triton, Kubernetes, Inference, Sandboxes