38 model serving engineer jobs at 33 companies in Tiburon, CA

1mo
Save
Mark Applied
Hide
Inference Infrastructure Engineer, Serving
Palo Alto, California, United States
$275k-$475k/yr OnsiteFull Time
Elorian AI
Elorian AI: AI lab building multimodal models for advanced visual reasoning.
3+ YOE3+ years building low-latency, high-throughput inference serving systems; knowledge of quantization, batching, speculative decoding, KV cache; experience with vLLM/TensorRT-LLM/Triton/SGLang; multi-GPU model parallelism; C++/CUDA/Python; autoscaling and GPU cost optimization.
vLLM, TensorRT-LLM, Triton, SGLang, C++, CUDA, Python
3w
Save
Mark Applied
Hide
Senior Performance Co-Design Engineer, LLM Serving
Sunnyvale, California, United States
$174k-$252k/yr OnsiteFull Time
Google
GoogleNASDAQ: GOOGL: Provides online search, advertising, cloud computing, and consumer electronics.
5+ YOEBachelor's in CS/EE/CE or equivalent,5+ years in performance modeling/engineering or architecture,proficiency with C++ or Python,experience with ML serving and hardware/software co-design preferred.
C++, Python, TPU, Vertex AI
2mo
Save
Mark Applied
Hide
Senior Machine Learning Engineer, Ad Serving
New York or San Jose
$195k-$408k/yr HybridFull Time
Roku
RokuNASDAQ: ROKU: Operates a TV streaming platform and sells streaming hardware.
10+ YOE10+ years applying machine learning and optimization to production systems; deep statistics/ML expertise; production ML lifecycle experience (feature engineering, model serving, monitoring); strong software skills in Python, SQL, Java/Scala; excellent communication.
Python, SQL, Java/Scala
1mo
Save
Mark Applied
Hide
Staff Software Engineer, Foundation Model API
San Francisco or Mountain View
$190k-$265k/yr OnsiteFull Time
Databricks
Databricks: A unified platform for data analytics and artificial intelligence.
8+ YOE8+ years backend/infrastructure experience, distributed systems, scalable APIs, cloud-native infra, real-time serving/ML infra/GPU orchestration, strong Scala/Go/Python skills, product ownership and customer engagement.
OpenAI, Anthropic, Gemini, Qwen, GPT-OSS, Llama, SageMaker, Vertex AI, Azure ML, FMAPI, Apache Spark™, Delta Lake, MLflow, Scala, Go, Python
1w
Save
Mark Applied
Hide
Senior Software Engineer, Data & Model
Fremont, California, United States
$150k-$300k/yr OnsiteFull Time
Dexmate
Dexmate: Building AI-powered mobile humanoid robots for industrial automation.
5+ YOERequires 5+ years in software, data or ML infrastructure, or distributed systems; experience with data pipelines, distributed computing, model serving, cloud infrastructure, containers, orchestration, and ML workflows.
AI, Physical AI, cloud infrastructure, storage, containers, orchestration
2d
Save
Mark Applied
Hide
Inference Performance Engineer
San Francisco, California, United States
HybridFull Time
Adaption
Adaption: Develops efficient AI systems that adapt in real-time.
5+ YOE5+ years in ML systems, inference infrastructure, or performance engineering; model-serving expertise; Python and systems-language proficiency; and GPU performance experience with measurable cost or latency improvements.
vLLM, SGLang, TensorRT-LLM, Python, C++, Rust, CUDA, NCCL
1mo
Save
Mark Applied
Hide
Staff, MLOps Engineer
New York City or San Francisco or United States
$220k-$280k/yr RemoteFull Time
Sequen AI: AI-native ranking engine for enterprise search and recommendations.
4+ YOEMinimum 4+ years MLOps or ML/platform engineering; expertise with low-latency model serving, Python and PyTorch; cloud (AWS/GCP/Azure), Docker, Kubernetes, MLflow; strong distributed systems and pipeline experience.
Python, PyTorch, AWS, GCP, Azure, Docker, Kubernetes, MLflow, Rust, vLLM, Triton
3mo
Save
Mark Applied
Hide
Machine Learning Engineer, Inference & Serving (Speech LLM) - San Francisco
San Francisco, California, United States
$180k-$270k/yr HybridFull Time
Plaud
Plaud: Develops AI-powered voice recorders and automated transcription software.
Experience building and deploying high-throughput, ultra-low-latency inference for LLMs or speech models; optimize latency/throughput; manage KV cache; understand GPU memory hierarchies; collaborate across ML and backend teams.
vLLM, TensorRT-LLM, SGLang, NVIDIA Triton Inference Server, WebSockets, WebRTC, CUDA, PTQ, FP8, INT8, AWQ, GPTQ, Tensor Parallelism, Kubernetes
2mo
Save
Mark Applied
Hide
Infrastructure Engineer
Redwood City, California, United States
HybridFull Time
Vantaca
Vantaca: AI software for community association and HOA management.
8+ YOE8+ years in infrastructure/DevOps/SRE; strong cloud expertise; experience with CI/CD, PostgreSQL, Redis, APM, model serving, vector databases, GPU optimization, and LLM deployment.
PostgreSQL, Redis, APM, CI/CD, vector databases, model serving frameworks, LLM
2mo
Save
Mark Applied
Hide
AI Field Engineer - AI Natives
San Mateo, California, United States
FieldFull Time
Fireworks AI
Fireworks AI: Provides high-performance generative AI model inference and deployment infrastructure.
5+ YOE5+ years in customer-facing technical engineering roles, strong Python and Kubernetes skills, experience with LLM inference, model serving and fine-tuning, cloud GPU deployment across major clouds, and exceptional communication.
Python, Kubernetes, vLLM, SGLang, TensorRT-LLM, AWS, Microsoft Azure, GCP, Azure AI Foundry, AWS Bedrock, SageMaker, GCP Vertex
2mo
Save
Mark Applied
Hide
Staff Machine Learning Engineer
San Francisco or Minneapolis
HybridFull Time
Shipt
Shipt: Provides same-day delivery services from local retailers via app.
5+ YOE5+ years of machine learning and backend software engineering; backend in Go/Java and Python; embeddings, similarity search, ranking models; ML pipelines; distributed systems; SQL/NoSQL; API serving; A/B testing.
Go, Java, Python, MLflow, Kubeflow, Airflow, REST, gRPC, Model servers, SQL, NoSQL
2mo
Save
Mark Applied
Hide
Senior Machine Learning Engineer
Palo Alto, California, United States
$189k-$283k/yr OnsiteFull Time
Rubrik
RubrikNYSE: RBRK: Secures enterprise data across cloud and on-premises environments.
2+ YOEBachelor's in a technical field required, 2+ years production ML experience, proficiency in Python and PyTorch, experience training/fine-tuning/distilling language models, serving low-latency models, and building closed-loop data and evaluation pipelines.
Python, PyTorch, vLLM, SGLang, TensorRT-LLM, LoRA, DPO, RLAIF, RLHF, GRPO, FP8, INT8, KV-cache, MCP, LiteLLM, Google ADK, Azure AI Foundry, Vertex AI
1mo
Save
Mark Applied
Hide
Machine Learning Engineer (Staff)
San Francisco or Menlo Park
$220k-$270k/yr HybridFull Time
Sprinter Health
Sprinter Health: Mobile provider of in-home diagnostic and preventive healthcare services.
8+ YOE8+ years building production ML systems and infrastructure; experience with training/serving pipelines, feature pipelines, monitoring, deployment, cloud, containers, CI/CD, and model governance.
CI/CD, APIs, containers, feature stores, MLOps, LLM
6d
Save
Mark Applied
Hide
ML Infrastructure Engineer
San Mateo, California, United States
OnsiteFull Time
Clera
Clera: AI talent agent matching professionals with high-growth startup roles
5+ YOERequires 5+ years building production ML inference or model-serving systems, experience scaling for latency and reliability, Docker, Kubernetes, distributed systems, observability tools, cloud platforms, and Python, Go, Rust, C++, or Java.
TensorFlow Serving, TorchServe, Triton, KServe, Docker, Kubernetes, Prometheus, Grafana, ELK stack, AWS, GCP, Azure, Python, Go, Rust, C++, Java, Neo4j, Amazon Neptune
2mo
Save
Mark Applied
Hide
Machine Learning Engineer 5 - Decisioning & Optimization
New York or Los Angeles or Los Gatos or Seattle
$466k-$750k/yr OnsiteFull Time
Netflix
NetflixNASDAQ: NFLX: Global video streaming and media production service.
7+ YOE7+ years software engineering; 3+ years ML infrastructure, model serving, or ML platform experience in ads/real-time decisioning; real-time model serving with sub-20ms latency; proficiency in Java, Python, or Scala; experience with ML serving frameworks and real-time feature pipelines; strong model monitoring and production readiness.
Java, Python, Scala, ML serving frameworks, feature stores, model registries
3mo
Save
Mark Applied
Hide
Staff Machine Learning Engineer (Pricing)
San Francisco, California, United States
HybridFull Time
GoFundMe
GoFundMe: Provides an online platform for personal and nonprofit fundraising.
7+ YOE7+ years building production ML systems; Python and ML libraries; pricing/monetization or growth optimization; real-time model serving; data engineering; ML monitoring; leadership.
Python, PyTorch, TensorFlow, Scikit-learn, AWS, Databricks, Docker, Kubernetes, FastAPI, Terraform, Snowflake, GitHub
2mo
Save
Mark Applied
Hide
Head of Developer Relations
San Mateo, California, United States
$160k-$240k/yr HybridFull Time
Parasail
Parasail: Provides scalable cloud infrastructure for AI model inference.
6+ YOE6+ years combined software engineering, DevRel, and technical content experience; strong technical fluency in inference/model serving; prior experience growing developer communities and producing technical content and benchmarks.
Discord, Hacker News
2mo
Save
Mark Applied
Hide
Member of Technical Staff, Developer Relations
San Francisco, California, United States
$200k-$400k/yr OnsiteFull Time
Inferact
Inferact: An AI infrastructure building vLLM to accelerate and scale model inference.
Bachelor's or equivalent experience in CS/engineering/ML, deep knowledge of LLM inference and model serving, experience with vLLM or adjacent systems, strong technical writing/teaching portfolio, and ability to build demos and tutorials.
vLLM, SGLang, TensorRT-LLM, TGI, LoRAX, Ray Serve, FlashInfer, BentoML, Baseten, CUDA, PyTorch, Modal, Predibase, Together AI, Anyscale, LMSYS
2mo
Save
Mark Applied
Hide
Sr Machine Learning Services Engineer
San Jose, California, United States
$152k-$265k/yr OnsiteFull Time
Adobe
AdobeNASDAQ: ADBE: Provides software for digital media creation and marketing analytics
5+ YOE5+ years building and operating ML systems with GPU workloads, designing large-scale cloud services, optimizing model inference, and collaborating with research and engineering teams; proficiency with Python, PyTorch/TensorFlow, model serving tools, containerization, orchestration, and AWS.
Python, PyTorch, TensorFlow, NVIDIA Triton, TorchServe, ONNX, AIT, AOT, CUDA, Docker, Kubernetes, AWS
2d
Save
Mark Applied
Hide
Principal AI Engineer
Austin or Menlo Park or Reston
HybridFull Time
Seekr
Seekr: Transparent AI platform for enterprise and government decision-making.
Production AI/ML platform experience, strong Python or Go skills, Docker, Kubernetes, Linux, APIs, infrastructure-as-code, model serving, security, testing, observability, and edge deployment expertise.
Python, Go, Docker, Kubernetes, Linux, APIs, ONNX, TensorRT, PyTorch, Triton, vLLM