42 model serving engineer jobs at 33 companies in San Bruno, CA
1mo
Save
Mark Applied
Hide
1mo
Inference Infrastructure Engineer, Serving
Palo Alto, California, United States
$275k-$475k/yrOnsiteFull Time
Elorian: AI research lab building multimodal visual-reasoning models for machines, robotics teams, engineers, and scientific organizations.
3+ YOE3+ years building low-latency, high-throughput inference serving systems; knowledge of quantization, batching, speculative decoding, KV cache; experience with vLLM/TensorRT-LLM/Triton/SGLang; multi-GPU model parallelism; C++/CUDA/Python; autoscaling and GPU cost optimization.
GoogleNASDAQ: GOOG, GOOGL: Global technology specializing in internet-related services and products.
5+ YOEBachelor's in CS/EE/CE or equivalent,5+ years in performance modeling/engineering or architecture,proficiency with C++ or Python,experience with ML serving and hardware/software co-design preferred.
RokuNASDAQ: ROKU: TV streaming platform powering the global television ecosystem.
10+ YOE10+ years applying machine learning and optimization to production systems; deep statistics/ML expertise; production ML lifecycle experience (feature engineering, model serving, monitoring); strong software skills in Python, SQL, Java/Scala; excellent communication.
Distributed Systems Engineer 5 - Core Ad Serving Platform
New York City or Seattle or Los Angeles or Los Gatos
$388k-$619k/yrOnsiteFull Time
NetflixNASDAQ: NFLX: Global subscription-based streaming entertainment service and content producer.
7+ YOE7+ years experience with at least 4+ years in Ads domain, expertise building and operating large-scale distributed systems, ad-server components, API and data model design, SLO-driven development, and incident response.
Dexmate: Robotics and AI building dexterous mobile humanoid robots for industrial automation.
5+ YOERequires 5+ years in software, data or ML infrastructure, or distributed systems; experience with data pipelines, distributed computing, model serving, cloud infrastructure, containers, orchestration, and ML workflows.
OpenAI: AI research and deployment focused on beneficial AGI.
Strong systems programming in C++, Rust, or Python; experience with runtimes, distributed systems, compilers, kernels, or serving infrastructure; knowledge of LLM inference and hardware-software performance optimization.
Adaption: AI building adaptive intelligence that continually learns for industries, languages, and specialized workflows.
5+ YOE5+ years in ML systems, inference infrastructure, or performance engineering; model-serving expertise; Python and systems-language proficiency; and GPU performance experience with measurable cost or latency improvements.
Sequen AI: AI personalization and ranking platform serving enterprise consumer companies with dynamic search, recommendations, and discovery.
4+ YOEMinimum 4+ years MLOps or ML/platform engineering; expertise with low-latency model serving, Python and PyTorch; cloud (AWS/GCP/Azure), Docker, Kubernetes, MLflow; strong distributed systems and pipeline experience.
NVIDIANASDAQ: NVDA: Computing platform for AI and accelerated graphics.
16+ YOE7+ MgmtRequires 16+ years in technical roles, 7–10+ years of leadership, and BS/MS or equivalent experience in systems or software engineering. Expertise in AI infrastructure, distributed systems, GPU architecture, CUDA, kernels, and cloud-native model serving.
CUDA, Dynamo, TensorRT-LLM, vLLM, SGLang, Linux, Kubernetes, Ray
Machine Learning Engineer, Inference & Serving (Speech LLM) - San Francisco
San Francisco, California, United States
$180k-$270k/yrHybridFull Time
Plaud: AI note-taking hardware and software serving professionals with voice recording, transcription, and meeting-summary tools.
Experience building and deploying high-throughput, ultra-low-latency inference for LLMs or speech models; optimize latency/throughput; manage KV cache; understand GPU memory hierarchies; collaborate across ML and backend teams.
HOAi: AI-first community association management software serving management companies, vendors, boards, and homeowners.
8+ YOE8+ years in infrastructure/DevOps/SRE; strong cloud expertise; experience with CI/CD, PostgreSQL, Redis, APM, model serving, vector databases, GPU optimization, and LLM deployment.
PostgreSQL, Redis, APM, CI/CD, vector databases, model serving frameworks, LLM
Fireworks AI: AI is a private AI infrastructure serving developers and enterprises with model training and inference.
5+ YOE5+ years in customer-facing technical engineering roles, strong Python and Kubernetes skills, experience with LLM inference, model serving and fine-tuning, cloud GPU deployment across major clouds, and exceptional communication.
Python, Kubernetes, vLLM, SGLang, TensorRT-LLM, AWS, Microsoft Azure, GCP, Azure AI Foundry, AWS Bedrock, SageMaker, GCP Vertex
Shipt: Target-owned retail-tech providing same-day grocery and household-essential delivery to U.S. consumers.
5+ YOE5+ years of machine learning and backend software engineering; backend in Go/Java and Python; embeddings, similarity search, ranking models; ML pipelines; distributed systems; SQL/NoSQL; API serving; A/B testing.
RubrikNYSE: RBRK: Public cybersecurity and AI operations software helping organizations protect, monitor, and recover data, identities, and workloads.
2+ YOEBachelor's in a technical field required, 2+ years production ML experience, proficiency in Python and PyTorch, experience training/fine-tuning/distilling language models, serving low-latency models, and building closed-loop data and evaluation pipelines.
Python, PyTorch, vLLM, SGLang, TensorRT-LLM, LoRA, DPO, RLAIF, RLHF, GRPO, FP8, INT8, KV-cache, MCP, LiteLLM, Google ADK, Azure AI Foundry, Vertex AI
Sprinter Health: Private mobile healthcare provider delivering in-home diagnostics and preventive care to patients through nurses and virtual clinicians.
8+ YOE8+ years building production ML systems and infrastructure; experience with training/serving pipelines, feature pipelines, monitoring, deployment, cloud, containers, CI/CD, and model governance.
Clera: AI-powered talent agent matching candidates to startup roles.
5+ YOE5+ years building production ML inference or model-serving systems. Requires scalable distributed systems, Docker, Kubernetes, cloud experience, observability tooling, and proficiency in Python, Go, Rust, C++, or Java.
NetflixNASDAQ: NFLX: Global subscription-based streaming entertainment service and content producer.
7+ YOE7+ years software engineering; 3+ years ML infrastructure, model serving, or ML platform experience in ads/real-time decisioning; real-time model serving with sub-20ms latency; proficiency in Java, Python, or Scala; experience with ML serving frameworks and real-time feature pipelines; strong model monitoring and production readiness.
Java, Python, Scala, ML serving frameworks, feature stores, model registries
GoFundMe: For-profit crowdfunding platform helping people and nonprofits raise money for personal, charitable, and community causes.
7+ YOE7+ years building production ML systems; Python and ML libraries; pricing/monetization or growth optimization; real-time model serving; data engineering; ML monitoring; leadership.