20 model serving engineer jobs at 13 companies in Marina, CA
🚀PromotedHiringCafe
ML Engineer - Inference & Model Deployment
Cupertino, CA, US
$250k-$310k/yrOn-SiteFull Time
HiringCafe: Building a 100× better job search engine to take on Indeed and LinkedIn.
Turn powerful AI and ML models into fast, reliable production systems. Own inference latency, throughput, model-serving architecture, multi-GPU systems, and production deployment for millions of users.
RokuNASDAQ: ROKU: Operates a TV streaming platform and sells streaming hardware.
10+ YOE10+ years applying machine learning and optimization to production systems; deep statistics/ML expertise; production ML lifecycle experience (feature engineering, model serving, monitoring); strong software skills in Python, SQL, Java/Scala; excellent communication.
Senior Machine Learning Engineer, Model Serving Infrastructure (Multiple Positions)
San Jose, California, United States
$265k-$388k/yrOnsiteFull Time
ByteDance: Developing AI-driven content platforms and mobile applications.
2+ YOEMaster's (plus 2 years) or Bachelor's (plus 5 years) in a quantitative field; 2+ years coding in Python or C++; Linux development experience; ML, system design, and production deployment experience.
Distributed Systems Engineer 5 - Decisioning & Optimization
New York or Los Angeles or Los Gatos or Seattle
$388k-$619k/yrOnsiteFull Time
NetflixNASDAQ: NFLX: Global video streaming and media production service.
7+ YOE7+ years building distributed systems; ads domain experience; ML model serving; backend APIs and ad tech systems; low-latency serving; collaboration across teams.
ML model serving, real-time inference, APIs, ad servers, bidders, pacing, routing, calibration serving
NVIDIANASDAQ: NVDA: Designs graphics processing units and artificial intelligence hardware.
3+ YOE3+ years supporting live-site production environments, SRE on-call experience, strong Python/Java scripting, Kubernetes and AWS expertise, experience with ML model serving and MLOps frameworks.
Distributed Systems Engineer 6 - Decisioning & Optimization
New York City or Seattle or Los Angeles or Los Gatos
$499k-$900k/yrOnsiteFull Time
NetflixNASDAQ: NFLX: Provider of global streaming entertainment and video content.
10+ YOE10+ years building large-scale distributed systems (3+ years in ads); expertise in ML model serving with sub-20ms P99, ad serving/bidding/pacing systems, API and platform design, technical leadership, and cross-functional collaboration.
AdobeNASDAQ: ADBE: Provides software for digital media creation and marketing analytics
5+ YOE5+ years building and operating ML systems with GPU workloads, designing large-scale cloud services, optimizing model inference, and collaborating with research and engineering teams; proficiency with Python, PyTorch/TensorFlow, model serving tools, containerization, orchestration, and AWS.
WalmartNYSE: WMT: Multinational retail operating discount stores and supermarkets.
10+ YOE10+ years building and operating large-scale distributed systems; deep AI/ML systems and model-serving experience; cloud-native platform and observability expertise; strong communication and technical leadership.
KlearNow.AI: AI platform for digital customs brokerage and logistics management.
Engineering leader with hands-on DevSecOps experience across AWS and GCP, CI/CD and IaC expertise, AI infrastructure (GPU/compute provisioning, model serving), cloud security, and scripting in Python/Bash/Go/Java.
Sr. Staff Software Engineer, Systems Infrastructure
Sunnyvale or Mountain View
$198k-$326k/yrHybridFull Time
LinkedInNASDAQ: MSFT: Professional social network for career development and job recruitment.
8+ YOEBA/BS or equivalent; 8+ years engineering experience building large-scale ML/model serving and GPU systems; experience with CUDA, PyTorch/TensorFlow, C++/Go/Python/Java; strong systems and performance optimization skills.
AppleNASDAQ: AAPL: Designs and sells consumer electronics, software, and online services.
Expertise in model serving, deployment pipelines, distributed systems, and ML platform infrastructure to build and scale ML-powered features for content tagging, ranking, and personalization.
Sr. Staff AI Engineer - On-Prem AI Infrastructure & Agentic Systems
San Jose, California, United States
$140k-$165k/yrOnsiteFull Time
SK hynix Memory Solutions AmericaKorea Exchange: 000660: Develops semiconductor controllers and firmware for enterprise data storage.
2+ YOE2+ years in AI/ML engineering with on-prem/private-cloud deployment, experience building agentic AI and RAG pipelines, model fine-tuning (LoRA/QLoRA), Python/Linux/Docker/Kubernetes proficiency, and familiarity with vector DBs and AI serving frameworks.
5+ YOEBachelor's degree or equivalent experience, 5+ years Android development with Kotlin/Java, client-side performance optimization, EMR not applicable; experience with on-device model serving and privacy reviews preferred.
Spring or San Jose or Durham or Fort Collins or Andover
$170k-$413k/yrHybridFull Time
Hewlett Packard EnterpriseNYSE: HPE: Provides edge-to-cloud IT infrastructure and platform services.
15+ YOE5+ MgmtBachelor's in CS/engineering required; 15+ years product experience with 5+ years leading product teams in AI/ML or data infrastructure. Experience with AI platforms, model serving, GPU ecosystems, GTM and monetization.
Spring or San Jose or Durham or Fort Collins or Andover
$170k-$413k/yrHybridFull Time
Hewlett Packard EnterpriseNYSE: HPE: Provides global edge-to-cloud technology solutions and IT infrastructure services.
15+ YOE5+ MgmtBachelor's in CS/engineering required; 15+ years product management experience with 5+ years AI/ML product leadership; experience building AI/ML platforms, inference/model serving, GPU ecosystem, executive communication, and GTM strategy.
GreenLake, GenAI, GPU, agent frameworks, AI tools, MLOps, DevOps, SaaS
Livingston or New York or Sunnyvale or San Francisco or Bellevue
$207k-$275k/yrOnsiteFull Time
CoreWeaveNASDAQ: CRWV: Cloud platform providing GPU-accelerated infrastructure for AI workloads.
10+ YOE10+ years in distributed systems/ML infrastructure or production AI engineering; 5+ years with AI runtime systems; expertise in model serving, batching, execution isolation, GPU memory management, and Kubernetes; strong customer-facing and commercial skills.