Sequen AI: AI-native ranking engine for enterprise search and recommendations.
4+ YOEMinimum 4+ years MLOps or ML/platform engineering; expertise with low-latency model serving, Python and PyTorch; cloud (AWS/GCP/Azure), Docker, Kubernetes, MLflow; strong distributed systems and pipeline experience.
Machine Learning Engineer, Inference & Serving (Speech LLM) - San Francisco
San Francisco, California, United States
$180k-$270k/yrHybridFull Time
Plaud: Develops AI-powered voice recorders and automated transcription software.
Experience building and deploying high-throughput, ultra-low-latency inference for LLMs or speech models; optimize latency/throughput; manage KV cache; understand GPU memory hierarchies; collaborate across ML and backend teams.
Shipt: Provides same-day delivery services from local retailers via app.
5+ YOE5+ years of machine learning and backend software engineering; backend in Go/Java and Python; embeddings, similarity search, ranking models; ML pipelines; distributed systems; SQL/NoSQL; API serving; A/B testing.
Sprinter Health: Mobile provider of in-home diagnostic and preventive healthcare services.
8+ YOE8+ years building production ML systems and infrastructure; experience with training/serving pipelines, feature pipelines, monitoring, deployment, cloud, containers, CI/CD, and model governance.
Sprinter Health: Mobile provider of in-home diagnostic and preventive healthcare services.
Experience building production ML systems, training and inference pipelines, model serving, monitoring/observability, cloud and containers, strong Python and software engineering skills.
Gap Inc.NYSE: GAP: Global specialty retailer of apparel and accessories.
10+ YOE10+ years building production ML systems; strong Python and software engineering; experience with ML frameworks, model serving, MLOps, cloud platforms, and distributed data processing.
GoFundMe: Provides an online platform for personal and nonprofit fundraising.
7+ YOE7+ years building production ML systems; Python and ML libraries; pricing/monetization or growth optimization; real-time model serving; data engineering; ML monitoring; leadership.
Inferact: An AI infrastructure building vLLM to accelerate and scale model inference.
Bachelor's or equivalent experience in CS/engineering/ML, deep knowledge of LLM inference and model serving, experience with vLLM or adjacent systems, strong technical writing/teaching portfolio, and ability to build demos and tutorials.
vLLM, SGLang, TensorRT-LLM, TGI, LoRAX, Ray Serve, FlashInfer, BentoML, Baseten, CUDA, PyTorch, Modal, Predibase, Together AI, Anyscale, LMSYS
Together AI: Cloud platform for training and deploying artificial intelligence models.
8+ YOE8+ years ML engineering experience focused on model serving, inference optimization, and ML infrastructure at production scale; strong Python/PyTorch, GPU optimization, and system design skills; leadership and developer tooling experience; domain knowledge in speech/audio AI preferred.
United States or Canada or San Francisco or New York City or Seattle or Boston or Chicago or Denver or Austin or Portland
$195k-$358k/yrRemoteFull Time
Outschool: Marketplace for live online small-group classes for children.
10+ YOE5+ Mgmt10+ years in data/analytics/engineering roles, 5+ years managing data or analytics teams, strong SQL/dbt/data modeling skills, experience with Redshift, self-serve analytics tools, experimentation, and stakeholder communication.
Livingston or New York or Sunnyvale or San Francisco or Bellevue
$207k-$275k/yrOnsiteFull Time
CoreWeaveNASDAQ: CRWV: Cloud platform providing GPU-accelerated infrastructure for AI workloads.
10+ YOE10+ years in distributed systems/ML infrastructure or production AI engineering; 5+ years with AI runtime systems; expertise in model serving, batching, execution isolation, GPU memory management, and Kubernetes; strong customer-facing and commercial skills.