20 model serving engineer jobs at 13 companies in Marina, CA

PromotedHiringCafe
ML Engineer - Inference & Model Deployment
Cupertino, CA, US
$250k-$310k/yr On-SiteFull Time
HiringCafe
HiringCafe: Building a 100× better job search engine to take on Indeed and LinkedIn.
Turn powerful AI and ML models into fast, reliable production systems. Own inference latency, throughput, model-serving architecture, multi-GPU systems, and production deployment for millions of users.
Python, PyTorch, vLLM, SGLang, TensorRT, LLMs
1mo
Save
Mark Applied
Hide
Senior Machine Learning Engineer, Ad Serving
New York or San Jose
$195k-$408k/yr HybridFull Time
Roku
RokuNASDAQ: ROKU: Operates a TV streaming platform and sells streaming hardware.
10+ YOE10+ years applying machine learning and optimization to production systems; deep statistics/ML expertise; production ML lifecycle experience (feature engineering, model serving, monitoring); strong software skills in Python, SQL, Java/Scala; excellent communication.
Python, SQL, Java/Scala
1w
Save
Mark Applied
Hide
Senior Machine Learning Engineer, Model Serving Infrastructure (Multiple Positions)
San Jose, California, United States
$265k-$388k/yr OnsiteFull Time
ByteDance
ByteDance: Developing AI-driven content platforms and mobile applications.
2+ YOEMaster's (plus 2 years) or Bachelor's (plus 5 years) in a quantitative field; 2+ years coding in Python or C++; Linux development experience; ML, system design, and production deployment experience.
Python, C++, Linux
2mo
Save
Mark Applied
Hide
Distributed Systems Engineer 5 - Decisioning & Optimization
New York or Los Angeles or Los Gatos or Seattle
$388k-$619k/yr OnsiteFull Time
Netflix
NetflixNASDAQ: NFLX: Global video streaming and media production service.
7+ YOE7+ years building distributed systems; ads domain experience; ML model serving; backend APIs and ad tech systems; low-latency serving; collaboration across teams.
ML model serving, real-time inference, APIs, ad servers, bidders, pacing, routing, calibration serving
1w
Save
Mark Applied
Hide
Site Reliability Engineer, Cloud
Santa Clara, California, United States
$116k-$219k/yr OnsiteFull Time
NVIDIA
NVIDIANASDAQ: NVDA: Designs graphics processing units and artificial intelligence hardware.
3+ YOE3+ years supporting live-site production environments, SRE on-call experience, strong Python/Java scripting, Kubernetes and AWS expertise, experience with ML model serving and MLOps frameworks.
Python, Java, Kubernetes, MLflow, Kubeflow, Airflow, AWS SageMaker, AWS, FastAPI, Triton, TorchServe, Akamai Edge Redirector Cloudlets, Akamai Cloudlets Policy Manager, Akamai CDN
2mo
Save
Mark Applied
Hide
Distributed Systems Engineer 6 - Decisioning & Optimization
New York City or Seattle or Los Angeles or Los Gatos
$499k-$900k/yr OnsiteFull Time
Netflix
NetflixNASDAQ: NFLX: Provider of global streaming entertainment and video content.
10+ YOE10+ years building large-scale distributed systems (3+ years in ads); expertise in ML model serving with sub-20ms P99, ad serving/bidding/pacing systems, API and platform design, technical leadership, and cross-functional collaboration.
4w
Save
Mark Applied
Hide
Sr Machine Learning Services Engineer
San Jose, California, United States
$152k-$265k/yr OnsiteFull Time
Adobe
AdobeNASDAQ: ADBE: Provides software for digital media creation and marketing analytics
5+ YOE5+ years building and operating ML systems with GPU workloads, designing large-scale cloud services, optimizing model inference, and collaborating with research and engineering teams; proficiency with Python, PyTorch/TensorFlow, model serving tools, containerization, orchestration, and AWS.
Python, PyTorch, TensorFlow, NVIDIA Triton, TorchServe, ONNX, AIT, AOT, CUDA, Docker, Kubernetes, AWS
1mo
Save
Mark Applied
Hide
(USA) Distinguished, Software Engineer
Bentonville or Sunnyvale
$130k-$260k/yr OnsiteFull Time
Walmart
WalmartNYSE: WMT: Multinational retail operating discount stores and supermarkets.
10+ YOE10+ years building and operating large-scale distributed systems; deep AI/ML systems and model-serving experience; cloud-native platform and observability expertise; strong communication and technical leadership.
JSON Schema, OpenAPI, Pydantic, protobuf, WCAG 2.2, CI/CD
2w
Save
Mark Applied
Hide
DevOps, MLOps & Security Engineering Lead
San Jose, California, United States
$140k-$170k/yr HybridFull Time
KlearNow.AI
KlearNow.AI: AI platform for digital customs brokerage and logistics management.
Engineering leader with hands-on DevSecOps experience across AWS and GCP, CI/CD and IaC expertise, AI infrastructure (GPU/compute provisioning, model serving), cloud security, and scripting in Python/Bash/Go/Java.
AWS, GCP, SageMaker, Vertex AI, EKS, GKE, Python, Bash, Go, Java
3w
Save
Mark Applied
Hide
Sr. Staff Software Engineer, Systems Infrastructure
Sunnyvale or Mountain View
$198k-$326k/yr HybridFull Time
LinkedInNASDAQ: MSFT: Professional social network for career development and job recruitment.
8+ YOEBA/BS or equivalent; 8+ years engineering experience building large-scale ML/model serving and GPU systems; experience with CUDA, PyTorch/TensorFlow, C++/Go/Python/Java; strong systems and performance optimization skills.
SGLang, vLLM, Triton, TensorRT, Ray, XLA, TVM, PyTorch, TensorFlow, CUDA, C++, Go, Python, Java
2w
Save
Mark Applied
Hide
Sr. Machine Learning Engineer - Apple News
Cupertino, California, United States
OnsiteFull Time
Apple
AppleNASDAQ: AAPL: Designs and sells consumer electronics, software, and online services.
Expertise in model serving, deployment pipelines, distributed systems, and ML platform infrastructure to build and scale ML-powered features for content tagging, ranking, and personalization.
3w
Save
Mark Applied
Hide
Sr. Staff AI Engineer - On-Prem AI Infrastructure & Agentic Systems
San Jose, California, United States
$140k-$165k/yr OnsiteFull Time
SK hynix Memory Solutions America
SK hynix Memory Solutions AmericaKorea Exchange: 000660: Develops semiconductor controllers and firmware for enterprise data storage.
2+ YOE2+ years in AI/ML engineering with on-prem/private-cloud deployment, experience building agentic AI and RAG pipelines, model fine-tuning (LoRA/QLoRA), Python/Linux/Docker/Kubernetes proficiency, and familiarity with vector DBs and AI serving frameworks.
vLLM, TGI, Triton, Milvus, Qdrant, FAISS, Kubernetes, Helm, Docker, LangGraph, AutoGen, LoRA, QLoRA, Model Control Protocols (MCP), Python, Linux, Pinecone, Ollama, ONNX, TensorRT, GGUF, BabyAGI, LangSmith, Weights & Biases, Prometheus, Grafana, CI/CD
2w
Save
Mark Applied
Hide
Senior Software Engineer, Android Dialer, Calling Protection
San Jose, California, United States
$174k-$253k/yr OnsiteFull Time
Google
GoogleNASDAQ: GOOGL: Provides online search, advertising, cloud computing, and consumer electronics.
5+ YOEBachelor's degree or equivalent experience, 5+ years Android development with Kotlin/Java, client-side performance optimization, EMR not applicable; experience with on-device model serving and privacy reviews preferred.
Kotlin, Java, Android, Gemini, Gemini Nano, AICore
1mo
Save
Mark Applied
Hide
Director of Product Management – AI Essentials
Spring or San Jose or Durham or Fort Collins or Andover
$170k-$413k/yr HybridFull Time
Hewlett Packard Enterprise
Hewlett Packard EnterpriseNYSE: HPE: Provides edge-to-cloud IT infrastructure and platform services.
15+ YOE5+ MgmtBachelor's in CS/engineering required; 15+ years product experience with 5+ years leading product teams in AI/ML or data infrastructure. Experience with AI platforms, model serving, GPU ecosystems, GTM and monetization.
GenAI, GPU, MLOps, GreenLake, AI tools, AI/ML, SaaS, DevOps, agent frameworks
1mo
Save
Mark Applied
Hide
Director of Product Management – AI Essentials
Spring or San Jose or Durham or Fort Collins or Andover
$170k-$413k/yr HybridFull Time
Hewlett Packard Enterprise
Hewlett Packard EnterpriseNYSE: HPE: Provides global edge-to-cloud technology solutions and IT infrastructure services.
15+ YOE5+ MgmtBachelor's in CS/engineering required; 15+ years product management experience with 5+ years AI/ML product leadership; experience building AI/ML platforms, inference/model serving, GPU ecosystem, executive communication, and GTM strategy.
GreenLake, GenAI, GPU, agent frameworks, AI tools, MLOps, DevOps, SaaS
3w
Save
Mark Applied
Hide
Solution Specialist, AI Runtime Services
Livingston or New York or Sunnyvale or San Francisco or Bellevue
$207k-$275k/yr OnsiteFull Time
CoreWeave
CoreWeaveNASDAQ: CRWV: Cloud platform providing GPU-accelerated infrastructure for AI workloads.
10+ YOE10+ years in distributed systems/ML infrastructure or production AI engineering; 5+ years with AI runtime systems; expertise in model serving, batching, execution isolation, GPU memory management, and Kubernetes; strong customer-facing and commercial skills.
vLLM, TensorRT-LLM, TGI, Triton, Kubernetes, Inference, Sandboxes