94 model serving engineer jobs at 76 companies in United States
1mo
Save
Mark Applied
Hide
1mo
Inference Infrastructure Engineer, Serving
Palo Alto, California, United States
$275k-$475k/yrOnsiteFull Time
Elorian: AI research lab building multimodal visual-reasoning models for machines, robotics teams, engineers, and scientific organizations.
3+ YOE3+ years building low-latency, high-throughput inference serving systems; knowledge of quantization, batching, speculative decoding, KV cache; experience with vLLM/TensorRT-LLM/Triton/SGLang; multi-GPU model parallelism; C++/CUDA/Python; autoscaling and GPU cost optimization.
GoogleNASDAQ: GOOG, GOOGL: Global technology specializing in internet-related services and products.
5+ YOEBachelor's in CS/EE/CE or equivalent,5+ years in performance modeling/engineering or architecture,proficiency with C++ or Python,experience with ML serving and hardware/software co-design preferred.
Member of Technical Staff — Model Optimization and Inference
Seattle, Washington, United States
$250k-$350k/yrOnsiteFull Time
Nuance Labs: Private AI research building real-time audiovisual foundation models for face-to-face conversational AI.
Deep expertise in LLM and diffusion-model inference optimization, KV cache strategies, quantization (INT8/INT4, GPTQ/AWQ), profiling/benchmarking, and strong Python/PyTorch skills; familiarity with CUDA/Triton and inference-serving frameworks.
RokuNASDAQ: ROKU: TV streaming platform powering the global television ecosystem.
10+ YOE10+ years applying machine learning and optimization to production systems; deep statistics/ML expertise; production ML lifecycle experience (feature engineering, model serving, monitoring); strong software skills in Python, SQL, Java/Scala; excellent communication.
NVIDIANASDAQ: NVDA: Computing platform for AI and accelerated graphics.
7+ YOEBachelor’s degree in computer science, engineering, or equivalent experience; 7+ years developing LLM serving infrastructure; experience with model customization, quantization, fine-tuning, platform APIs, and cross-functional collaboration.
Distributed Systems Engineer 5 - Core Ad Serving Platform
New York City or Seattle or Los Angeles or Los Gatos
$388k-$619k/yrOnsiteFull Time
NetflixNASDAQ: NFLX: Global subscription-based streaming entertainment service and content producer.
7+ YOE7+ years experience with at least 4+ years in Ads domain, expertise building and operating large-scale distributed systems, ad-server components, API and data model design, SLO-driven development, and incident response.
Dexmate: Robotics and AI building dexterous mobile humanoid robots for industrial automation.
5+ YOERequires 5+ years in software, data or ML infrastructure, or distributed systems; experience with data pipelines, distributed computing, model serving, cloud infrastructure, containers, orchestration, and ML workflows.
OpenAI: AI research and deployment focused on beneficial AGI.
Strong systems programming in C++, Rust, or Python; experience with runtimes, distributed systems, compilers, kernels, or serving infrastructure; knowledge of LLM inference and hardware-software performance optimization.
Fundamental: AI serving enterprises and governments with foundation models for tabular data prediction.
5+ YOE5+ years MLOps/DevOps experience, degree in CS/Engineering or equivalent, experience with model serving, Kubernetes, cloud providers, IaC, Python/Bash/Go, and observability tools.
Adaption: AI building adaptive intelligence that continually learns for industries, languages, and specialized workflows.
5+ YOE5+ years in ML systems, inference infrastructure, or performance engineering; model-serving expertise; Python and systems-language proficiency; and GPU performance experience with measurable cost or latency improvements.
Sequen AI: AI personalization and ranking platform serving enterprise consumer companies with dynamic search, recommendations, and discovery.
4+ YOEMinimum 4+ years MLOps or ML/platform engineering; expertise with low-latency model serving, Python and PyTorch; cloud (AWS/GCP/Azure), Docker, Kubernetes, MLflow; strong distributed systems and pipeline experience.
Machine Learning Engineer, Inference & Serving (Speech LLM) - San Francisco
San Francisco, California, United States
$180k-$270k/yrHybridFull Time
Plaud: AI note-taking hardware and software serving professionals with voice recording, transcription, and meeting-summary tools.
Experience building and deploying high-throughput, ultra-low-latency inference for LLMs or speech models; optimize latency/throughput; manage KV cache; understand GPU memory hierarchies; collaborate across ML and backend teams.
HOAi: AI-first community association management software serving management companies, vendors, boards, and homeowners.
8+ YOE8+ years in infrastructure/DevOps/SRE; strong cloud expertise; experience with CI/CD, PostgreSQL, Redis, APM, model serving, vector databases, GPU optimization, and LLM deployment.
PostgreSQL, Redis, APM, CI/CD, vector databases, model serving frameworks, LLM
Fireworks AI: AI is a private AI infrastructure serving developers and enterprises with model training and inference.
5+ YOE5+ years in customer-facing technical engineering roles, strong Python and Kubernetes skills, experience with LLM inference, model serving and fine-tuning, cloud GPU deployment across major clouds, and exceptional communication.
Python, Kubernetes, vLLM, SGLang, TensorRT-LLM, AWS, Microsoft Azure, GCP, Azure AI Foundry, AWS Bedrock, SageMaker, GCP Vertex
Seekr Technologies: Private American enterprise AI providing explainable, secure AI software and hardware to government and critical-infrastructure customers.
5+ YOE5–8 years building distributed systems or cloud/platform services; production Kubernetes and ML infra experience; strong software engineering in Python and Go/Rust/C++; familiarity with GPU inference, model serving frameworks, and cloud platforms.
Shipt: Target-owned retail-tech providing same-day grocery and household-essential delivery to U.S. consumers.
5+ YOE5+ years of machine learning and backend software engineering; backend in Go/Java and Python; embeddings, similarity search, ranking models; ML pipelines; distributed systems; SQL/NoSQL; API serving; A/B testing.
AtlassianNASDAQ: TEAM: Unleashing the potential of every team.
Design, develop, and deploy production ML systems for search and retrieval; ensure low-latency, high-availability serving; integrate models with Triton and PyTorch; drive observability, cost optimization, and mentor engineers.
RubrikNYSE: RBRK: Public cybersecurity and AI operations software helping organizations protect, monitor, and recover data, identities, and workloads.
2+ YOEBachelor's in a technical field required, 2+ years production ML experience, proficiency in Python and PyTorch, experience training/fine-tuning/distilling language models, serving low-latency models, and building closed-loop data and evaluation pipelines.
Python, PyTorch, vLLM, SGLang, TensorRT-LLM, LoRA, DPO, RLAIF, RLHF, GRPO, FP8, INT8, KV-cache, MCP, LiteLLM, Google ADK, Azure AI Foundry, Vertex AI