42 model serving engineer jobs at 33 companies in San Bruno, CA

1mo
Save
Mark Applied
Hide
Inference Infrastructure Engineer, Serving
Palo Alto, California, United States
$275k-$475k/yr OnsiteFull Time
Elorian
Elorian: AI research lab building multimodal visual-reasoning models for machines, robotics teams, engineers, and scientific organizations.
3+ YOE3+ years building low-latency, high-throughput inference serving systems; knowledge of quantization, batching, speculative decoding, KV cache; experience with vLLM/TensorRT-LLM/Triton/SGLang; multi-GPU model parallelism; C++/CUDA/Python; autoscaling and GPU cost optimization.
vLLM, TensorRT-LLM, Triton, SGLang, C++, CUDA, Python
4w
Save
Mark Applied
Hide
Senior Performance Co-Design Engineer, LLM Serving
Sunnyvale, California, United States
$174k-$252k/yr OnsiteFull Time
Google
GoogleNASDAQ: GOOG, GOOGL: Global technology specializing in internet-related services and products.
5+ YOEBachelor's in CS/EE/CE or equivalent,5+ years in performance modeling/engineering or architecture,proficiency with C++ or Python,experience with ML serving and hardware/software co-design preferred.
C++, Python, TPU, Vertex AI
2mo
Save
Mark Applied
Hide
Senior Machine Learning Engineer, Ad Serving
New York or San Jose
$195k-$408k/yr HybridFull Time
Roku
RokuNASDAQ: ROKU: TV streaming platform powering the global television ecosystem.
10+ YOE10+ years applying machine learning and optimization to production systems; deep statistics/ML expertise; production ML lifecycle experience (feature engineering, model serving, monitoring); strong software skills in Python, SQL, Java/Scala; excellent communication.
Python, SQL, Java/Scala
1mo
Save
Mark Applied
Hide
Staff Software Engineer, Foundation Model API
San Francisco or Mountain View
$190k-$265k/yr OnsiteFull Time
Databricks
Databricks: Data and AI software providing a unified platform.
8+ YOE8+ years backend/infrastructure experience, distributed systems, scalable APIs, cloud-native infra, real-time serving/ML infra/GPU orchestration, strong Scala/Go/Python skills, product ownership and customer engagement.
OpenAI, Anthropic, Gemini, Qwen, GPT-OSS, Llama, SageMaker, Vertex AI, Azure ML, FMAPI, Apache Spark™, Delta Lake, MLflow, Scala, Go, Python
1mo
Save
Mark Applied
Hide
Distributed Systems Engineer 5 - Core Ad Serving Platform
New York City or Seattle or Los Angeles or Los Gatos
$388k-$619k/yr OnsiteFull Time
Netflix
NetflixNASDAQ: NFLX: Global subscription-based streaming entertainment service and content producer.
7+ YOE7+ years experience with at least 4+ years in Ads domain, expertise building and operating large-scale distributed systems, ad-server components, API and data model design, SLO-driven development, and incident response.
2w
Save
Mark Applied
Hide
Senior Software Engineer, Data & Model
Fremont, California, United States
$150k-$300k/yr OnsiteFull Time
Dexmate
Dexmate: Robotics and AI building dexterous mobile humanoid robots for industrial automation.
5+ YOERequires 5+ years in software, data or ML infrastructure, or distributed systems; experience with data pipelines, distributed computing, model serving, cloud infrastructure, containers, orchestration, and ML workflows.
AI, Physical AI, cloud infrastructure, storage, containers, orchestration
4d
Save
Mark Applied
Hide
Software Engineer, Model Runtime
San Francisco, California, United States
$266k-$445k/yr HybridFull Time
OpenAI
OpenAI: AI research and deployment focused on beneficial AGI.
Strong systems programming in C++, Rust, or Python; experience with runtimes, distributed systems, compilers, kernels, or serving infrastructure; knowledge of LLM inference and hardware-software performance optimization.
C++, Rust, Python, vLLM, SGLang
1w
Save
Mark Applied
Hide
Inference Performance Engineer
San Francisco, California, United States
HybridFull Time
Adaption
Adaption: AI building adaptive intelligence that continually learns for industries, languages, and specialized workflows.
5+ YOE5+ years in ML systems, inference infrastructure, or performance engineering; model-serving expertise; Python and systems-language proficiency; and GPU performance experience with measurable cost or latency improvements.
vLLM, SGLang, TensorRT-LLM, Python, C++, Rust, CUDA, NCCL
1mo
Save
Mark Applied
Hide
Staff, MLOps Engineer
New York City or San Francisco or United States
$220k-$280k/yr RemoteFull Time
Sequen AI
Sequen AI: AI personalization and ranking platform serving enterprise consumer companies with dynamic search, recommendations, and discovery.
4+ YOEMinimum 4+ years MLOps or ML/platform engineering; expertise with low-latency model serving, Python and PyTorch; cloud (AWS/GCP/Azure), Docker, Kubernetes, MLflow; strong distributed systems and pipeline experience.
Python, PyTorch, AWS, GCP, Azure, Docker, Kubernetes, MLflow, Rust, vLLM, Triton
2w
Save
Mark Applied
Hide
Distinguished Engineer, Scaled Out Inferencing
Santa Clara or California
$320k-$489k/yr HybridFull Time
NVIDIA
NVIDIANASDAQ: NVDA: Computing platform for AI and accelerated graphics.
16+ YOE7+ MgmtRequires 16+ years in technical roles, 7–10+ years of leadership, and BS/MS or equivalent experience in systems or software engineering. Expertise in AI infrastructure, distributed systems, GPU architecture, CUDA, kernels, and cloud-native model serving.
CUDA, Dynamo, TensorRT-LLM, vLLM, SGLang, Linux, Kubernetes, Ray
3mo
Save
Mark Applied
Hide
Machine Learning Engineer, Inference & Serving (Speech LLM) - San Francisco
San Francisco, California, United States
$180k-$270k/yr HybridFull Time
Plaud
Plaud: AI note-taking hardware and software serving professionals with voice recording, transcription, and meeting-summary tools.
Experience building and deploying high-throughput, ultra-low-latency inference for LLMs or speech models; optimize latency/throughput; manage KV cache; understand GPU memory hierarchies; collaborate across ML and backend teams.
vLLM, TensorRT-LLM, SGLang, NVIDIA Triton Inference Server, WebSockets, WebRTC, CUDA, PTQ, FP8, INT8, AWQ, GPTQ, Tensor Parallelism, Kubernetes
2mo
Save
Mark Applied
Hide
Infrastructure Engineer
Redwood City, California, United States
HybridFull Time
HOAi
HOAi: AI-first community association management software serving management companies, vendors, boards, and homeowners.
8+ YOE8+ years in infrastructure/DevOps/SRE; strong cloud expertise; experience with CI/CD, PostgreSQL, Redis, APM, model serving, vector databases, GPU optimization, and LLM deployment.
PostgreSQL, Redis, APM, CI/CD, vector databases, model serving frameworks, LLM
2mo
Save
Mark Applied
Hide
AI Field Engineer - AI Natives
San Mateo, California, United States
FieldFull Time
Fireworks AI
Fireworks AI: AI is a private AI infrastructure serving developers and enterprises with model training and inference.
5+ YOE5+ years in customer-facing technical engineering roles, strong Python and Kubernetes skills, experience with LLM inference, model serving and fine-tuning, cloud GPU deployment across major clouds, and exceptional communication.
Python, Kubernetes, vLLM, SGLang, TensorRT-LLM, AWS, Microsoft Azure, GCP, Azure AI Foundry, AWS Bedrock, SageMaker, GCP Vertex
3mo
Save
Mark Applied
Hide
Staff Machine Learning Engineer
San Francisco or Minneapolis
HybridFull Time
Shipt
Shipt: Target-owned retail-tech providing same-day grocery and household-essential delivery to U.S. consumers.
5+ YOE5+ years of machine learning and backend software engineering; backend in Go/Java and Python; embeddings, similarity search, ranking models; ML pipelines; distributed systems; SQL/NoSQL; API serving; A/B testing.
Go, Java, Python, MLflow, Kubeflow, Airflow, REST, gRPC, Model servers, SQL, NoSQL
2mo
Save
Mark Applied
Hide
Senior Machine Learning Engineer
Palo Alto, California, United States
$189k-$283k/yr OnsiteFull Time
Rubrik
RubrikNYSE: RBRK: Public cybersecurity and AI operations software helping organizations protect, monitor, and recover data, identities, and workloads.
2+ YOEBachelor's in a technical field required, 2+ years production ML experience, proficiency in Python and PyTorch, experience training/fine-tuning/distilling language models, serving low-latency models, and building closed-loop data and evaluation pipelines.
Python, PyTorch, vLLM, SGLang, TensorRT-LLM, LoRA, DPO, RLAIF, RLHF, GRPO, FP8, INT8, KV-cache, MCP, LiteLLM, Google ADK, Azure AI Foundry, Vertex AI
1mo
Save
Mark Applied
Hide
Machine Learning Engineer (Staff)
San Francisco or Menlo Park
$220k-$270k/yr HybridFull Time
Sprinter Health
Sprinter Health: Private mobile healthcare provider delivering in-home diagnostics and preventive care to patients through nurses and virtual clinicians.
8+ YOE8+ years building production ML systems and infrastructure; experience with training/serving pipelines, feature pipelines, monitoring, deployment, cloud, containers, CI/CD, and model governance.
CI/CD, APIs, containers, feature stores, MLOps, LLM
3d
Save
Mark Applied
Hide
ML Infrastructure Engineer
San Mateo, California, United States
OnsiteFull Time
Clera
Clera: AI-powered talent agent matching candidates to startup roles.
5+ YOE5+ years building production ML inference or model-serving systems. Requires scalable distributed systems, Docker, Kubernetes, cloud experience, observability tooling, and proficiency in Python, Go, Rust, C++, or Java.
TensorFlow Serving, TorchServe, Triton, KServe, Docker, Kubernetes, Prometheus, Grafana, ELK, AWS, GCP, Azure, Python, Go, Rust, C++, Java, Neo4j, Amazon Neptune
2mo
Save
Mark Applied
Hide
Machine Learning Engineer 5 - Decisioning & Optimization
New York or Los Angeles or Los Gatos or Seattle
$466k-$750k/yr OnsiteFull Time
Netflix
NetflixNASDAQ: NFLX: Global subscription-based streaming entertainment service and content producer.
7+ YOE7+ years software engineering; 3+ years ML infrastructure, model serving, or ML platform experience in ads/real-time decisioning; real-time model serving with sub-20ms latency; proficiency in Java, Python, or Scala; experience with ML serving frameworks and real-time feature pipelines; strong model monitoring and production readiness.
Java, Python, Scala, ML serving frameworks, feature stores, model registries
3mo
Save
Mark Applied
Hide
Staff Machine Learning Engineer (Pricing)
San Francisco, California, United States
HybridFull Time
GoFundMe
GoFundMe: For-profit crowdfunding platform helping people and nonprofits raise money for personal, charitable, and community causes.
7+ YOE7+ years building production ML systems; Python and ML libraries; pricing/monetization or growth optimization; real-time model serving; data engineering; ML monitoring; leadership.
Python, PyTorch, TensorFlow, Scikit-learn, AWS, Databricks, Docker, Kubernetes, FastAPI, Terraform, Snowflake, GitHub
2mo
Save
Mark Applied
Hide
Head of Developer Relations
San Mateo, California, United States
$160k-$240k/yr HybridFull Time
Parasail
Parasail: AI inference cloud providing production-ready open-model services for AI-native startups.
6+ YOE6+ years combined software engineering, DevRel, and technical content experience; strong technical fluency in inference/model serving; prior experience growing developer communities and producing technical content and benchmarks.
Discord, Hacker News