95 model serving engineer jobs at 76 companies in United States

1mo
Save
Mark Applied
Hide
Inference Infrastructure Engineer, Serving
Palo Alto, California, United States
$275k-$475k/yr OnsiteFull Time
Elorian
Elorian: AI research lab building multimodal visual-reasoning models for machines, robotics teams, engineers, and scientific organizations.
3+ YOE3+ years building low-latency, high-throughput inference serving systems; knowledge of quantization, batching, speculative decoding, KV cache; experience with vLLM/TensorRT-LLM/Triton/SGLang; multi-GPU model parallelism; C++/CUDA/Python; autoscaling and GPU cost optimization.
vLLM, TensorRT-LLM, Triton, SGLang, C++, CUDA, Python
1mo
Save
Mark Applied
Hide
Senior Performance Co-Design Engineer, LLM Serving
Sunnyvale, California, United States
$174k-$252k/yr OnsiteFull Time
Google
GoogleNASDAQ: GOOG, GOOGL: Global technology specializing in internet-related services and products.
5+ YOEBachelor's in CS/EE/CE or equivalent,5+ years in performance modeling/engineering or architecture,proficiency with C++ or Python,experience with ML serving and hardware/software co-design preferred.
C++, Python, TPU, Vertex AI
2mo
Save
Mark Applied
Hide
Member of Technical Staff — Model Optimization and Inference
Seattle, Washington, United States
$250k-$350k/yr OnsiteFull Time
Nuance Labs
Nuance Labs: Private AI research building real-time audiovisual foundation models for face-to-face conversational AI.
Deep expertise in LLM and diffusion-model inference optimization, KV cache strategies, quantization (INT8/INT4, GPTQ/AWQ), profiling/benchmarking, and strong Python/PyTorch skills; familiarity with CUDA/Triton and inference-serving frameworks.
vLLM, SGLang, TensorRT-LLM, Python, PyTorch, CUDA, Triton
2mo
Save
Mark Applied
Hide
Senior Machine Learning Engineer, Ad Serving
New York or San Jose
$195k-$408k/yr HybridFull Time
Roku
RokuNASDAQ: ROKU: TV streaming platform powering the global television ecosystem.
10+ YOE10+ years applying machine learning and optimization to production systems; deep statistics/ML expertise; production ML lifecycle experience (feature engineering, model serving, monitoring); strong software skills in Python, SQL, Java/Scala; excellent communication.
Python, SQL, Java/Scala
2w
Save
Mark Applied
Hide
Machine Learning Engineer, Foundation Model Services
Seattle, Washington, United States
OnsiteFull Time
Apple
AppleNASDAQ: AAPL: Designing and manufacturing consumer electronics, software, and digital services.
Build frameworks, services, and tools for production foundation models, optimizing and serving large language, vision, and speech models at scale.
5d
Save
Mark Applied
Hide
Senior Software Engineer, NIM Model Customization
Santa Clara, California, United States
$184k-$357k/yr OnsiteFull Time
NVIDIA
NVIDIANASDAQ: NVDA: Computing platform for AI and accelerated graphics.
7+ YOEBachelor’s degree in computer science, engineering, or equivalent experience; 7+ years developing LLM serving infrastructure; experience with model customization, quantization, fine-tuning, platform APIs, and cross-functional collaboration.
LoRA, QLoRA, supervised fine-tuning (SFT), reinforcement learning (RL), preference optimization, DPO, retrieval-augmented generation (RAG), vLLM, TRT-LLM
1mo
Save
Mark Applied
Hide
Staff Software Engineer, Foundation Model API
San Francisco or Mountain View
$190k-$265k/yr OnsiteFull Time
Databricks
Databricks: Data and AI software providing a unified platform.
8+ YOE8+ years backend/infrastructure experience, distributed systems, scalable APIs, cloud-native infra, real-time serving/ML infra/GPU orchestration, strong Scala/Go/Python skills, product ownership and customer engagement.
OpenAI, Anthropic, Gemini, Qwen, GPT-OSS, Llama, SageMaker, Vertex AI, Azure ML, FMAPI, Apache Spark™, Delta Lake, MLflow, Scala, Go, Python
1mo
Save
Mark Applied
Hide
Distributed Systems Engineer 5 - Core Ad Serving Platform
New York City or Seattle or Los Angeles or Los Gatos
$388k-$619k/yr OnsiteFull Time
Netflix
NetflixNASDAQ: NFLX: Global subscription-based streaming entertainment service and content producer.
7+ YOE7+ years experience with at least 4+ years in Ads domain, expertise building and operating large-scale distributed systems, ad-server components, API and data model design, SLO-driven development, and incident response.
2w
Save
Mark Applied
Hide
Senior Software Engineer, Data & Model
Fremont, California, United States
$150k-$300k/yr OnsiteFull Time
Dexmate
Dexmate: Robotics and AI building dexterous mobile humanoid robots for industrial automation.
5+ YOERequires 5+ years in software, data or ML infrastructure, or distributed systems; experience with data pipelines, distributed computing, model serving, cloud infrastructure, containers, orchestration, and ML workflows.
AI, Physical AI, cloud infrastructure, storage, containers, orchestration
1w
Save
Mark Applied
Hide
Software Engineer, Model Runtime
San Francisco, California, United States
$266k-$445k/yr HybridFull Time
OpenAI
OpenAI: AI research and deployment focused on beneficial AGI.
Strong systems programming in C++, Rust, or Python; experience with runtimes, distributed systems, compilers, kernels, or serving infrastructure; knowledge of LLM inference and hardware-software performance optimization.
C++, Rust, Python, vLLM, SGLang
1mo
Save
Mark Applied
Hide
MLOps Engineer
Europe or Israel or United States
RemoteFull Time
Fundamental
Fundamental: AI serving enterprises and governments with foundation models for tabular data prediction.
5+ YOE5+ years MLOps/DevOps experience, degree in CS/Engineering or equivalent, experience with model serving, Kubernetes, cloud providers, IaC, Python/Bash/Go, and observability tools.
TorchServe, TensorFlow, Triton, MLflow, WandB, PyTorch, TensorFlow Serving, KServe, Kubernetes, AWS, GCP, Azure, Terraform, Helm, GitOps, Python, Bash, Go, Prometheus, Grafana, Datadog, OpenTelemetry, FastAPI, Databricks, Snowflake
1w
Save
Mark Applied
Hide
Inference Performance Engineer
San Francisco, California, United States
HybridFull Time
Adaption
Adaption: AI building adaptive intelligence that continually learns for industries, languages, and specialized workflows.
5+ YOE5+ years in ML systems, inference infrastructure, or performance engineering; model-serving expertise; Python and systems-language proficiency; and GPU performance experience with measurable cost or latency improvements.
vLLM, SGLang, TensorRT-LLM, Python, C++, Rust, CUDA, NCCL
1mo
Save
Mark Applied
Hide
Staff, MLOps Engineer
New York City or San Francisco or United States
$220k-$280k/yr RemoteFull Time
Sequen AI
Sequen AI: AI personalization and ranking platform serving enterprise consumer companies with dynamic search, recommendations, and discovery.
4+ YOEMinimum 4+ years MLOps or ML/platform engineering; expertise with low-latency model serving, Python and PyTorch; cloud (AWS/GCP/Azure), Docker, Kubernetes, MLflow; strong distributed systems and pipeline experience.
Python, PyTorch, AWS, GCP, Azure, Docker, Kubernetes, MLflow, Rust, vLLM, Triton
3mo
Save
Mark Applied
Hide
Machine Learning Engineer, Inference & Serving (Speech LLM) - San Francisco
San Francisco, California, United States
$180k-$270k/yr HybridFull Time
Plaud
Plaud: AI note-taking hardware and software serving professionals with voice recording, transcription, and meeting-summary tools.
Experience building and deploying high-throughput, ultra-low-latency inference for LLMs or speech models; optimize latency/throughput; manage KV cache; understand GPU memory hierarchies; collaborate across ML and backend teams.
vLLM, TensorRT-LLM, SGLang, NVIDIA Triton Inference Server, WebSockets, WebRTC, CUDA, PTQ, FP8, INT8, AWQ, GPTQ, Tensor Parallelism, Kubernetes
2mo
Save
Mark Applied
Hide
Infrastructure Engineer
Redwood City, California, United States
HybridFull Time
HOAi
HOAi: AI-first community association management software serving management companies, vendors, boards, and homeowners.
8+ YOE8+ years in infrastructure/DevOps/SRE; strong cloud expertise; experience with CI/CD, PostgreSQL, Redis, APM, model serving, vector databases, GPU optimization, and LLM deployment.
PostgreSQL, Redis, APM, CI/CD, vector databases, model serving frameworks, LLM
2mo
Save
Mark Applied
Hide
AI Field Engineer - AI Natives
San Mateo, California, United States
FieldFull Time
Fireworks AI
Fireworks AI: AI is a private AI infrastructure serving developers and enterprises with model training and inference.
5+ YOE5+ years in customer-facing technical engineering roles, strong Python and Kubernetes skills, experience with LLM inference, model serving and fine-tuning, cloud GPU deployment across major clouds, and exceptional communication.
Python, Kubernetes, vLLM, SGLang, TensorRT-LLM, AWS, Microsoft Azure, GCP, Azure AI Foundry, AWS Bedrock, SageMaker, GCP Vertex
1mo
Save
Mark Applied
Hide
Senior AI Infrastructure Engineer
Austin or Reston or United States
HybridFull Time
Seekr Technologies
Seekr Technologies: Private American enterprise AI providing explainable, secure AI software and hardware to government and critical-infrastructure customers.
5+ YOE5–8 years building distributed systems or cloud/platform services; production Kubernetes and ML infra experience; strong software engineering in Python and Go/Rust/C++; familiarity with GPU inference, model serving frameworks, and cloud platforms.
Kubernetes, Helm, Argo CD, Docker, Prometheus, Grafana, OpenTelemetry, Python, Go, Rust, C++, vLLM, SGLang, TensorRT-LLM, Triton Inference Server, Ray Serve, AWS, Azure, Oracle Cloud Infrastructure, Google Cloud Platform
3mo
Save
Mark Applied
Hide
Staff Machine Learning Engineer
San Francisco or Minneapolis
HybridFull Time
Shipt
Shipt: Target-owned retail-tech providing same-day grocery and household-essential delivery to U.S. consumers.
5+ YOE5+ years of machine learning and backend software engineering; backend in Go/Java and Python; embeddings, similarity search, ranking models; ML pipelines; distributed systems; SQL/NoSQL; API serving; A/B testing.
Go, Java, Python, MLflow, Kubeflow, Airflow, REST, gRPC, Model servers, SQL, NoSQL
1mo
Save
Mark Applied
Hide
Senior Machine Learning System Engineer
Seattle, Washington, United States
$149k-$235k/yr HybridFull Time
Atlassian
AtlassianNASDAQ: TEAM: Unleashing the potential of every team.
Design, develop, and deploy production ML systems for search and retrieval; ensure low-latency, high-availability serving; integrate models with Triton and PyTorch; drive observability, cost optimization, and mentor engineers.
Triton, PyTorch
2mo
Save
Mark Applied
Hide
Senior Machine Learning Engineer
Palo Alto, California, United States
$189k-$283k/yr OnsiteFull Time
Rubrik
RubrikNYSE: RBRK: Public cybersecurity and AI operations software helping organizations protect, monitor, and recover data, identities, and workloads.
2+ YOEBachelor's in a technical field required, 2+ years production ML experience, proficiency in Python and PyTorch, experience training/fine-tuning/distilling language models, serving low-latency models, and building closed-loop data and evaluation pipelines.
Python, PyTorch, vLLM, SGLang, TensorRT-LLM, LoRA, DPO, RLAIF, RLHF, GRPO, FP8, INT8, KV-cache, MCP, LiteLLM, Google ADK, Azure AI Foundry, Vertex AI