40 model serving engineer jobs at 34 companies in San Jose, CA

1w
Save
Mark Applied
Hide
Inference Infrastructure Engineer, Serving
Palo Alto, California, United States
$275k-$475k/yr OnsiteFull Time
Elorian AI
Elorian AI: AI lab building multimodal models for advanced visual reasoning.
3+ YOE3+ years building low-latency, high-throughput inference serving systems; knowledge of quantization, batching, speculative decoding, KV cache; experience with vLLM/TensorRT-LLM/Triton/SGLang; multi-GPU model parallelism; C++/CUDA/Python; autoscaling and GPU cost optimization.
vLLM, TensorRT-LLM, Triton, SGLang, C++, CUDA, Python
3d
Save
Mark Applied
Hide
Senior Performance Co-Design Engineer, LLM Serving
Sunnyvale, California, United States
$174k-$252k/yr OnsiteFull Time
Google
GoogleNASDAQ: GOOGL: Provides online search, advertising, cloud computing, and consumer electronics.
5+ YOEBachelor's in CS/EE/CE or equivalent,5+ years in performance modeling/engineering or architecture,proficiency with C++ or Python,experience with ML serving and hardware/software co-design preferred.
C++, Python, TPU, Vertex AI
1mo
Save
Mark Applied
Hide
Senior Machine Learning Engineer, Ad Serving
New York or San Jose
$195k-$408k/yr HybridFull Time
Roku
RokuNASDAQ: ROKU: Operates a TV streaming platform and sells streaming hardware.
10+ YOE10+ years applying machine learning and optimization to production systems; deep statistics/ML expertise; production ML lifecycle experience (feature engineering, model serving, monitoring); strong software skills in Python, SQL, Java/Scala; excellent communication.
Python, SQL, Java/Scala
3w
Save
Mark Applied
Hide
Senior Machine Learning Engineer, Model Serving Infrastructure (Multiple Positions)
San Jose, California, United States
$265k-$388k/yr OnsiteFull Time
ByteDance
ByteDance: Developing AI-driven content platforms and mobile applications.
2+ YOEMaster's (plus 2 years) or Bachelor's (plus 5 years) in a quantitative field; 2+ years coding in Python or C++; Linux development experience; ML, system design, and production deployment experience.
Python, C++, Linux
5d
Save
Mark Applied
Hide
Distributed Systems Engineer 5 - Core Ad Serving Platform
New York City or Seattle or Los Angeles or Los Gatos
$388k-$619k/yr OnsiteFull Time
Netflix
NetflixNASDAQ: NFLX: Provider of global streaming entertainment and video content.
7+ YOE7+ years experience with at least 4+ years in Ads domain, expertise building and operating large-scale distributed systems, ad-server components, API and data model design, SLO-driven development, and incident response.
2w
Save
Mark Applied
Hide
Staff Software Engineer, Foundation Model API
San Francisco or Mountain View
$190k-$265k/yr OnsiteFull Time
Databricks
Databricks: A unified platform for data analytics and artificial intelligence.
8+ YOE8+ years backend/infrastructure experience, distributed systems, scalable APIs, cloud-native infra, real-time serving/ML infra/GPU orchestration, strong Scala/Go/Python skills, product ownership and customer engagement.
OpenAI, Anthropic, Gemini, Qwen, GPT-OSS, Llama, SageMaker, Vertex AI, Azure ML, FMAPI, Apache Spark™, Delta Lake, MLflow, Scala, Go, Python
3w
Save
Mark Applied
Hide
Staff, MLOps Engineer
New York City or San Francisco or United States
$220k-$280k/yr RemoteFull Time
Sequen AI: AI-native ranking engine for enterprise search and recommendations.
4+ YOEMinimum 4+ years MLOps or ML/platform engineering; expertise with low-latency model serving, Python and PyTorch; cloud (AWS/GCP/Azure), Docker, Kubernetes, MLflow; strong distributed systems and pipeline experience.
Python, PyTorch, AWS, GCP, Azure, Docker, Kubernetes, MLflow, Rust, vLLM, Triton
2mo
Save
Mark Applied
Hide
Machine Learning Engineer, Inference & Serving (Speech LLM) - San Francisco
San Francisco, California, United States
$180k-$270k/yr HybridFull Time
Plaud
Plaud: Develops AI-powered voice recorders and automated transcription software.
Experience building and deploying high-throughput, ultra-low-latency inference for LLMs or speech models; optimize latency/throughput; manage KV cache; understand GPU memory hierarchies; collaborate across ML and backend teams.
vLLM, TensorRT-LLM, SGLang, NVIDIA Triton Inference Server, WebSockets, WebRTC, CUDA, PTQ, FP8, INT8, AWQ, GPTQ, Tensor Parallelism, Kubernetes
1mo
Save
Mark Applied
Hide
Infrastructure Engineer
Redwood City, California, United States
HybridFull Time
Vantaca
Vantaca: AI software for community association and HOA management.
8+ YOE8+ years in infrastructure/DevOps/SRE; strong cloud expertise; experience with CI/CD, PostgreSQL, Redis, APM, model serving, vector databases, GPU optimization, and LLM deployment.
PostgreSQL, Redis, APM, CI/CD, vector databases, model serving frameworks, LLM
1mo
Save
Mark Applied
Hide
AI Field Engineer - AI Natives
San Mateo, California, United States
FieldFull Time
Fireworks AI
Fireworks AI: Provides high-performance generative AI model inference and deployment infrastructure.
5+ YOE5+ years in customer-facing technical engineering roles, strong Python and Kubernetes skills, experience with LLM inference, model serving and fine-tuning, cloud GPU deployment across major clouds, and exceptional communication.
Python, Kubernetes, vLLM, SGLang, TensorRT-LLM, AWS, Microsoft Azure, GCP, Azure AI Foundry, AWS Bedrock, SageMaker, GCP Vertex
2mo
Save
Mark Applied
Hide
Staff Machine Learning Engineer
San Francisco or Minneapolis
HybridFull Time
Shipt
Shipt: Provides same-day delivery services from local retailers via app.
5+ YOE5+ years of machine learning and backend software engineering; backend in Go/Java and Python; embeddings, similarity search, ranking models; ML pipelines; distributed systems; SQL/NoSQL; API serving; A/B testing.
Go, Java, Python, MLflow, Kubeflow, Airflow, REST, gRPC, Model servers, SQL, NoSQL
1mo
Save
Mark Applied
Hide
Senior Machine Learning Engineer
Palo Alto, California, United States
$189k-$283k/yr OnsiteFull Time
Rubrik
RubrikNYSE: RBRK: Secures enterprise data across cloud and on-premises environments.
2+ YOEBachelor's in a technical field required, 2+ years production ML experience, proficiency in Python and PyTorch, experience training/fine-tuning/distilling language models, serving low-latency models, and building closed-loop data and evaluation pipelines.
Python, PyTorch, vLLM, SGLang, TensorRT-LLM, LoRA, DPO, RLAIF, RLHF, GRPO, FP8, INT8, KV-cache, MCP, LiteLLM, Google ADK, Azure AI Foundry, Vertex AI
1w
Save
Mark Applied
Hide
Machine Learning Engineer (Staff)
San Francisco or Menlo Park
$220k-$270k/yr HybridFull Time
Sprinter Health
Sprinter Health: Mobile provider of in-home diagnostic and preventive healthcare services.
8+ YOE8+ years building production ML systems and infrastructure; experience with training/serving pipelines, feature pipelines, monitoring, deployment, cloud, containers, CI/CD, and model governance.
CI/CD, APIs, containers, feature stores, MLOps, LLM
3w
Save
Mark Applied
Hide
Analytics Engineer, Revenue
San Mateo or Provo or United States
$181k-$245k/yr RemoteFull Time
GC AI
GC AI: AI-powered legal platform for in-house legal teams.
5+ YOE5+ years in data engineering/analytics with strong SQL, data modeling, dbt and BI experience; builds dimensional models, KPIs, and self-serve analytics; cross-functional collaboration with Revenue and Finance.
SQL, dbt, Looker, Tableau, Sigma, BigQuery, Snowflake, Python, CRM, Monte Carlo, Great Expectations
2mo
Save
Mark Applied
Hide
Machine Learning Engineer 5 - Decisioning & Optimization
New York or Los Angeles or Los Gatos or Seattle
$466k-$750k/yr OnsiteFull Time
Netflix
NetflixNASDAQ: NFLX: Global video streaming and media production service.
7+ YOE7+ years software engineering; 3+ years ML infrastructure, model serving, or ML platform experience in ads/real-time decisioning; real-time model serving with sub-20ms latency; proficiency in Java, Python, or Scala; experience with ML serving frameworks and real-time feature pipelines; strong model monitoring and production readiness.
Java, Python, Scala, ML serving frameworks, feature stores, model registries
1mo
Save
Mark Applied
Hide
Sr. ML Engineer – ML & Applied AI
San Francisco, California, United States
$181k-$236k/yr OnsiteFull Time
Gap Inc.
Gap Inc.NYSE: GAP: Global specialty retailer of apparel and accessories.
10+ YOE10+ years building production ML systems; strong Python and software engineering; experience with ML frameworks, model serving, MLOps, cloud platforms, and distributed data processing.
Python, scikit-learn, XGBoost, PyTorch, TensorFlow, FastAPI, Flask, Docker, Kubernetes, GCP, AWS, Azure, Git, Spark, PySpark, Databricks
2mo
Save
Mark Applied
Hide
Staff Machine Learning Engineer (Pricing)
San Francisco, California, United States
HybridFull Time
GoFundMe
GoFundMe: Provides an online platform for personal and nonprofit fundraising.
7+ YOE7+ years building production ML systems; Python and ML libraries; pricing/monetization or growth optimization; real-time model serving; data engineering; ML monitoring; leadership.
Python, PyTorch, TensorFlow, Scikit-learn, AWS, Databricks, Docker, Kubernetes, FastAPI, Terraform, Snowflake, GitHub
1mo
Save
Mark Applied
Hide
Head of Developer Relations
San Mateo, California, United States
$160k-$240k/yr HybridFull Time
Parasail
Parasail: Provides scalable cloud infrastructure for AI model inference.
6+ YOE6+ years combined software engineering, DevRel, and technical content experience; strong technical fluency in inference/model serving; prior experience growing developer communities and producing technical content and benchmarks.
Discord, Hacker News
1mo
Save
Mark Applied
Hide
Software Development Engineer II (Planning and Forecasting)
Palo Alto, California, United States
$162k-$182k/yr OnsiteFull Time
Quince
Quince: Sells high-quality apparel and home goods at accessible prices.
2+ YOE2-4 years production software engineering experience; strong fundamentals in data structures and algorithms; experience with data pipelines, APIs, and model-serving; AI-native engineering practice; Bachelor's in CS/Engineering or equivalent experience.
Google Meet, Zoom
1mo
Save
Mark Applied
Hide
Member of Technical Staff, Developer Relations
San Francisco, California, United States
$200k-$400k/yr OnsiteFull Time
Inferact
Inferact: An AI infrastructure building vLLM to accelerate and scale model inference.
Bachelor's or equivalent experience in CS/engineering/ML, deep knowledge of LLM inference and model serving, experience with vLLM or adjacent systems, strong technical writing/teaching portfolio, and ability to build demos and tutorials.
vLLM, SGLang, TensorRT-LLM, TGI, LoRAX, Ray Serve, FlashInfer, BentoML, Baseten, CUDA, PyTorch, Modal, Predibase, Together AI, Anyscale, LMSYS