8 model serving engineer jobs at 6 companies in Langley, WA

2mo
Save
Mark Applied
Hide
Member of Technical Staff — Model Optimization and Inference
Seattle, Washington, United States
$250k-$350k/yr OnsiteFull Time
Nuance Labs
Nuance Labs: A building photorealistic, real-time AI avatars and full-duplex audiovisual systems.
Deep expertise in LLM and diffusion-model inference optimization, KV cache strategies, quantization (INT8/INT4, GPTQ/AWQ), profiling/benchmarking, and strong Python/PyTorch skills; familiarity with CUDA/Triton and inference-serving frameworks.
vLLM, SGLang, TensorRT-LLM, Python, PyTorch, CUDA, Triton
1w
Save
Mark Applied
Hide
Distributed Systems Engineer 5 - Core Ad Serving Platform
New York City or Seattle or Los Angeles or Los Gatos
$388k-$619k/yr OnsiteFull Time
Netflix
NetflixNASDAQ: NFLX: Provider of global streaming entertainment and video content.
7+ YOE7+ years experience with at least 4+ years in Ads domain, expertise building and operating large-scale distributed systems, ad-server components, API and data model design, SLO-driven development, and incident response.
6d
Save
Mark Applied
Hide
Senior Machine Learning System Engineer
Seattle, Washington, United States
$149k-$235k/yr HybridFull Time
Atlassian
AtlassianNASDAQ: TEAM: Develops software for team collaboration and project management.
Design, develop, and deploy production ML systems for search and retrieval; ensure low-latency, high-availability serving; integrate models with Triton and PyTorch; drive observability, cost optimization, and mentor engineers.
Triton, PyTorch
2mo
Save
Mark Applied
Hide
Machine Learning Engineer 5 - Decisioning & Optimization
New York City or Seattle or Los Angeles or Los Gatos
$466k-$750k/yr OnsiteFull Time
Netflix
NetflixNASDAQ: NFLX: Provider of global streaming entertainment and video content.
7+ YOE7+ years software engineering experience with 3+ years on ML infrastructure or model serving; proficiency in Java, Python, or Scala; experience building high‑QPS, low‑latency model serving, feature serving, and model monitoring.
Java, Python, Scala, Chronon, Signal Service, JVM
2mo
Save
Mark Applied
Hide
Machine Learning Engineer 5 - Decisioning & Optimization
New York or Los Angeles or Los Gatos or Seattle
$466k-$750k/yr OnsiteFull Time
Netflix
NetflixNASDAQ: NFLX: Global video streaming and media production service.
7+ YOE7+ years software engineering; 3+ years ML infrastructure, model serving, or ML platform experience in ads/real-time decisioning; real-time model serving with sub-20ms latency; proficiency in Java, Python, or Scala; experience with ML serving frameworks and real-time feature pipelines; strong model monitoring and production readiness.
Java, Python, Scala, ML serving frameworks, feature stores, model registries
2mo
Save
Mark Applied
Hide
Staff Software Engineer, Machine Learning Platform
Seattle or South San Francisco
$224k-$336k/yr OnsiteFull Time
Stripe
Stripe: Provides global financial infrastructure and payment processing for businesses.
10+ YOE10+ years software development experience, technical leadership on large distributed systems and ML platform work, experience with model training/serving/orchestration, strong communication and cross-functional collaboration.
AWS, SageMaker, Bedrock, Databricks, OpenAI
1mo
Save
Mark Applied
Hide
Director of Data
United States or Canada or San Francisco or New York City or Seattle or Boston or Chicago or Denver or Austin or Portland
$195k-$358k/yr RemoteFull Time
Outschool
Outschool: Marketplace for live online small-group classes for children.
10+ YOE5+ Mgmt10+ years in data/analytics/engineering roles, 5+ years managing data or analytics teams, strong SQL/dbt/data modeling skills, experience with Redshift, self-serve analytics tools, experimentation, and stakeholder communication.
Redshift, dbt, Opensearch, Omni, Amplitude, Statsig, Looker, Mode, Hex, SQL, Covey Scout
1mo
Save
Mark Applied
Hide
Solution Specialist, AI Runtime Services
Livingston or New York or Sunnyvale or San Francisco or Bellevue
$207k-$275k/yr OnsiteFull Time
CoreWeave
CoreWeaveNASDAQ: CRWV: Cloud platform providing GPU-accelerated infrastructure for AI workloads.
10+ YOE10+ years in distributed systems/ML infrastructure or production AI engineering; 5+ years with AI runtime systems; expertise in model serving, batching, execution isolation, GPU memory management, and Kubernetes; strong customer-facing and commercial skills.
vLLM, TensorRT-LLM, TGI, Triton, Kubernetes, Inference, Sandboxes