13 model serving engineer jobs at 10 companies in Santa Rosa, CA

2w
Save
Mark Applied
Hide
Staff Software Engineer, Foundation Model API
San Francisco or Mountain View
$190k-$265k/yr OnsiteFull Time
Databricks
Databricks: A unified platform for data analytics and artificial intelligence.
8+ YOE8+ years backend/infrastructure experience, distributed systems, scalable APIs, cloud-native infra, real-time serving/ML infra/GPU orchestration, strong Scala/Go/Python skills, product ownership and customer engagement.
OpenAI, Anthropic, Gemini, Qwen, GPT-OSS, Llama, SageMaker, Vertex AI, Azure ML, FMAPI, Apache Spark™, Delta Lake, MLflow, Scala, Go, Python
1w
Save
Mark Applied
Hide
Staff Software Engineer- Foundation Model Inference
San Francisco or Mountain View
$190k-$265k/yr OnsiteFull Time
Databricks
Databricks: A unified platform for data analytics and artificial intelligence.
8+ YOE8+ years backend or infrastructure engineering experience; distributed systems, scalable APIs, real-time serving or ML/GPU orchestration experience; familiarity with service-oriented architecture, deployment pipelines, and observability.
OpenAI, Anthropic, Gemini, Qwen, GPT-OSS, Llama, SageMaker, Vertex AI, Azure ML, MLflow, PyTorch, Ray, vLLM, SGLang, Apache Spark, Delta Lake
3w
Save
Mark Applied
Hide
Staff, MLOps Engineer
New York City or San Francisco or United States
$220k-$280k/yr RemoteFull Time
Sequen AI: AI-native ranking engine for enterprise search and recommendations.
4+ YOEMinimum 4+ years MLOps or ML/platform engineering; expertise with low-latency model serving, Python and PyTorch; cloud (AWS/GCP/Azure), Docker, Kubernetes, MLflow; strong distributed systems and pipeline experience.
Python, PyTorch, AWS, GCP, Azure, Docker, Kubernetes, MLflow, Rust, vLLM, Triton
2mo
Save
Mark Applied
Hide
Machine Learning Engineer, Inference & Serving (Speech LLM) - San Francisco
San Francisco, California, United States
$180k-$270k/yr HybridFull Time
Plaud
Plaud: Develops AI-powered voice recorders and automated transcription software.
Experience building and deploying high-throughput, ultra-low-latency inference for LLMs or speech models; optimize latency/throughput; manage KV cache; understand GPU memory hierarchies; collaborate across ML and backend teams.
vLLM, TensorRT-LLM, SGLang, NVIDIA Triton Inference Server, WebSockets, WebRTC, CUDA, PTQ, FP8, INT8, AWQ, GPTQ, Tensor Parallelism, Kubernetes
2mo
Save
Mark Applied
Hide
Staff Machine Learning Engineer
San Francisco or Minneapolis
HybridFull Time
Shipt
Shipt: Provides same-day delivery services from local retailers via app.
5+ YOE5+ years of machine learning and backend software engineering; backend in Go/Java and Python; embeddings, similarity search, ranking models; ML pipelines; distributed systems; SQL/NoSQL; API serving; A/B testing.
Go, Java, Python, MLflow, Kubeflow, Airflow, REST, gRPC, Model servers, SQL, NoSQL
2w
Save
Mark Applied
Hide
Machine Learning Engineer (Staff)
San Francisco or Menlo Park
$220k-$270k/yr HybridFull Time
Sprinter Health
Sprinter Health: Mobile provider of in-home diagnostic and preventive healthcare services.
8+ YOE8+ years building production ML systems and infrastructure; experience with training/serving pipelines, feature pipelines, monitoring, deployment, cloud, containers, CI/CD, and model governance.
CI/CD, APIs, containers, feature stores, MLOps, LLM
1w
Save
Mark Applied
Hide
Machine Learning Engineer
San Francisco, California, United States
$140k-$200k/yr HybridFull Time
Sprinter Health
Sprinter Health: Mobile provider of in-home diagnostic and preventive healthcare services.
Experience building production ML systems, training and inference pipelines, model serving, monitoring/observability, cloud and containers, strong Python and software engineering skills.
Python, CI/CD, containers, APIs, feature stores
1mo
Save
Mark Applied
Hide
Sr. ML Engineer – ML & Applied AI
San Francisco, California, United States
$181k-$236k/yr OnsiteFull Time
Gap Inc.
Gap Inc.NYSE: GAP: Global specialty retailer of apparel and accessories.
10+ YOE10+ years building production ML systems; strong Python and software engineering; experience with ML frameworks, model serving, MLOps, cloud platforms, and distributed data processing.
Python, scikit-learn, XGBoost, PyTorch, TensorFlow, FastAPI, Flask, Docker, Kubernetes, GCP, AWS, Azure, Git, Spark, PySpark, Databricks
2mo
Save
Mark Applied
Hide
Staff Machine Learning Engineer (Pricing)
San Francisco, California, United States
HybridFull Time
GoFundMe
GoFundMe: Provides an online platform for personal and nonprofit fundraising.
7+ YOE7+ years building production ML systems; Python and ML libraries; pricing/monetization or growth optimization; real-time model serving; data engineering; ML monitoring; leadership.
Python, PyTorch, TensorFlow, Scikit-learn, AWS, Databricks, Docker, Kubernetes, FastAPI, Terraform, Snowflake, GitHub
1mo
Save
Mark Applied
Hide
Member of Technical Staff, Developer Relations
San Francisco, California, United States
$200k-$400k/yr OnsiteFull Time
Inferact
Inferact: An AI infrastructure building vLLM to accelerate and scale model inference.
Bachelor's or equivalent experience in CS/engineering/ML, deep knowledge of LLM inference and model serving, experience with vLLM or adjacent systems, strong technical writing/teaching portfolio, and ability to build demos and tutorials.
vLLM, SGLang, TensorRT-LLM, TGI, LoRAX, Ray Serve, FlashInfer, BentoML, Baseten, CUDA, PyTorch, Modal, Predibase, Together AI, Anyscale, LMSYS
2mo
Save
Mark Applied
Hide
Staff Machine Learning Engineer, Voice AI
San Francisco, California, United States
$220k-$280k/yr OnsiteFull Time
Together AI
Together AI: Cloud platform for training and deploying artificial intelligence models.
8+ YOE8+ years ML engineering experience focused on model serving, inference optimization, and ML infrastructure at production scale; strong Python/PyTorch, GPU optimization, and system design skills; leadership and developer tooling experience; domain knowledge in speech/audio AI preferred.
Python, PyTorch, CUDA, TensorRT-LLM, vLLM, SGLang
1mo
Save
Mark Applied
Hide
Director of Data
United States or Canada or San Francisco or New York City or Seattle or Boston or Chicago or Denver or Austin or Portland
$195k-$358k/yr RemoteFull Time
Outschool
Outschool: Marketplace for live online small-group classes for children.
10+ YOE5+ Mgmt10+ years in data/analytics/engineering roles, 5+ years managing data or analytics teams, strong SQL/dbt/data modeling skills, experience with Redshift, self-serve analytics tools, experimentation, and stakeholder communication.
Redshift, dbt, Opensearch, Omni, Amplitude, Statsig, Looker, Mode, Hex, SQL, Covey Scout
1mo
Save
Mark Applied
Hide
Solution Specialist, AI Runtime Services
Livingston or New York or Sunnyvale or San Francisco or Bellevue
$207k-$275k/yr OnsiteFull Time
CoreWeave
CoreWeaveNASDAQ: CRWV: Cloud platform providing GPU-accelerated infrastructure for AI workloads.
10+ YOE10+ years in distributed systems/ML infrastructure or production AI engineering; 5+ years with AI runtime systems; expertise in model serving, batching, execution isolation, GPU memory management, and Kubernetes; strong customer-facing and commercial skills.
vLLM, TensorRT-LLM, TGI, Triton, Kubernetes, Inference, Sandboxes