26 model serving engineer jobs at 21 companies in Aromas, CA
2w
Save
Mark Applied
Hide
2w
Inference Infrastructure Engineer, Serving
Palo Alto, California, United States
$275k-$475k/yrOnsiteFull Time
Elorian AI: AI lab building multimodal models for advanced visual reasoning.
3+ YOE3+ years building low-latency, high-throughput inference serving systems; knowledge of quantization, batching, speculative decoding, KV cache; experience with vLLM/TensorRT-LLM/Triton/SGLang; multi-GPU model parallelism; C++/CUDA/Python; autoscaling and GPU cost optimization.
5+ YOEBachelor's in CS/EE/CE or equivalent,5+ years in performance modeling/engineering or architecture,proficiency with C++ or Python,experience with ML serving and hardware/software co-design preferred.
RokuNASDAQ: ROKU: Operates a TV streaming platform and sells streaming hardware.
10+ YOE10+ years applying machine learning and optimization to production systems; deep statistics/ML expertise; production ML lifecycle experience (feature engineering, model serving, monitoring); strong software skills in Python, SQL, Java/Scala; excellent communication.
Senior Machine Learning Engineer, Model Serving Infrastructure (Multiple Positions)
San Jose, California, United States
$265k-$388k/yrOnsiteFull Time
ByteDance: Developing AI-driven content platforms and mobile applications.
2+ YOEMaster's (plus 2 years) or Bachelor's (plus 5 years) in a quantitative field; 2+ years coding in Python or C++; Linux development experience; ML, system design, and production deployment experience.
Distributed Systems Engineer 5 - Core Ad Serving Platform
New York City or Seattle or Los Angeles or Los Gatos
$388k-$619k/yrOnsiteFull Time
NetflixNASDAQ: NFLX: Provider of global streaming entertainment and video content.
7+ YOE7+ years experience with at least 4+ years in Ads domain, expertise building and operating large-scale distributed systems, ad-server components, API and data model design, SLO-driven development, and incident response.
RubrikNYSE: RBRK: Secures enterprise data across cloud and on-premises environments.
2+ YOEBachelor's in a technical field required, 2+ years production ML experience, proficiency in Python and PyTorch, experience training/fine-tuning/distilling language models, serving low-latency models, and building closed-loop data and evaluation pipelines.
Python, PyTorch, vLLM, SGLang, TensorRT-LLM, LoRA, DPO, RLAIF, RLHF, GRPO, FP8, INT8, KV-cache, MCP, LiteLLM, Google ADK, Azure AI Foundry, Vertex AI
Sprinter Health: Mobile provider of in-home diagnostic and preventive healthcare services.
8+ YOE8+ years building production ML systems and infrastructure; experience with training/serving pipelines, feature pipelines, monitoring, deployment, cloud, containers, CI/CD, and model governance.
NetflixNASDAQ: NFLX: Global video streaming and media production service.
7+ YOE7+ years software engineering; 3+ years ML infrastructure, model serving, or ML platform experience in ads/real-time decisioning; real-time model serving with sub-20ms latency; proficiency in Java, Python, or Scala; experience with ML serving frameworks and real-time feature pipelines; strong model monitoring and production readiness.
Java, Python, Scala, ML serving frameworks, feature stores, model registries
Software Development Engineer II (Planning and Forecasting)
Palo Alto, California, United States
$162k-$182k/yrOnsiteFull Time
Quince: Sells high-quality apparel and home goods at accessible prices.
2+ YOE2-4 years production software engineering experience; strong fundamentals in data structures and algorithms; experience with data pipelines, APIs, and model-serving; AI-native engineering practice; Bachelor's in CS/Engineering or equivalent experience.
AdobeNASDAQ: ADBE: Provides software for digital media creation and marketing analytics
5+ YOE5+ years building and operating ML systems with GPU workloads, designing large-scale cloud services, optimizing model inference, and collaborating with research and engineering teams; proficiency with Python, PyTorch/TensorFlow, model serving tools, containerization, orchestration, and AWS.
Siemens HealthineersXetra: SHL: Manufacturer of medical diagnostic and imaging equipment.
6+ YOE6+ years software development experience; designing and scaling distributed systems; backend proficiency (Python/Go/Java); strong systems architecture, observability, and deployment skills; ML model serving preferred.
SAPFrankfurt Stock Exchange: SAP: Sells enterprise resource planning and business management software solutions.
7+ YOEMaster's in Software Engineering, 7+ years development experience, 3+ years with SAP BTP & SAP AI Core, strong Java/Python/NodeJs skills, Docker/Kubernetes, ML model serving, PyTorch or TensorFlow, SQL/Postgres, CI/CD and cloud experience.
SAP BTP, SAP AI Core, SAP GenAI Hub, Java, Python, NodeJs, Docker, Kubernetes, PyTorch, TensorFlow, Git, SQL, Postgres, AWS, GCP, Azure, ReactJS, SAP UI5, CI/CD
WalmartNYSE: WMT: Multinational retail operating discount stores and supermarkets.
10+ YOE10+ years building and operating large-scale distributed systems; deep AI/ML systems and model-serving experience; cloud-native platform and observability expertise; strong communication and technical leadership.
KlearNow.AI: AI platform for digital customs brokerage and logistics management.
Engineering leader with hands-on DevSecOps experience across AWS and GCP, CI/CD and IaC expertise, AI infrastructure (GPU/compute provisioning, model serving), cloud security, and scripting in Python/Bash/Go/Java.
Sr. Staff Software Engineer, Systems Infrastructure
Sunnyvale or Mountain View
$198k-$326k/yrHybridFull Time
LinkedInNASDAQ: MSFT: Professional social network for career development and job recruitment.
8+ YOEBA/BS or equivalent; 8+ years engineering experience building large-scale ML/model serving and GPU systems; experience with CUDA, PyTorch/TensorFlow, C++/Go/Python/Java; strong systems and performance optimization skills.
Level AI: AI-native platform for contact center intelligence and automation
6+ YOE6+ years applying ML/AI in production, deep learning and LLM expertise, Python and PyTorch proficiency, experience training/serving large models, distributed training, strong communication and technical leadership.
AppleNASDAQ: AAPL: Designs and sells consumer electronics, software, and online services.
Expertise in model serving, deployment pipelines, distributed systems, and ML platform infrastructure to build and scale ML-powered features for content tagging, ranking, and personalization.
Sr. Staff AI Engineer - On-Prem AI Infrastructure & Agentic Systems
San Jose, California, United States
$140k-$165k/yrOnsiteFull Time
SK hynix Memory Solutions AmericaKorea Exchange: 000660: Develops semiconductor controllers and firmware for enterprise data storage.
2+ YOE2+ years in AI/ML engineering with on-prem/private-cloud deployment, experience building agentic AI and RAG pipelines, model fine-tuning (LoRA/QLoRA), Python/Linux/Docker/Kubernetes proficiency, and familiarity with vector DBs and AI serving frameworks.
Software Development Engineer, Sponsored Products and Brands
Palo Alto, California, United States
OnsiteFull Time
AmazonNASDAQ: AMZN: Global online retail and cloud computing technology provider.
5+ YOE5+ years professional software development, programming, system design/architecture experience, mentoring or tech lead experience; experience with ML infrastructure and model serving preferred.