39 llm serving engineer jobs at 36 companies in United States
1mo
Save
Mark Applied
Hide
1mo
Staff ML/LLM Ops Engineer
Seattle, Washington, United States
$213k-$272k/yrOnsiteFull Time
LiveView Technologies: Mobile surveillance and automated security deterrence systems.
8+ YOE8+ years engineering experience in ML infrastructure/MLOps; hands-on LLM/VLM ops; model CI/CD, serving, monitoring, guardrails; strong API and system design; BS/MS in CS/Engineering or equivalent.
LiveView Technologies: Provides mobile, solar-powered security units with AI-driven surveillance software.
8+ YOE8+ years in ML-infrastructure/MLOps building and operating model deployment, serving, CI/CD, monitoring, and LLM/VLM ops; strong API design and technical leadership; BS/MS in CS/Engineering or equivalent experience.
5+ YOEBachelor's in CS/EE/CE or equivalent,5+ years in performance modeling/engineering or architecture,proficiency with C++ or Python,experience with ML serving and hardware/software co-design preferred.
Elorian AI: AI lab building multimodal models for advanced visual reasoning.
3+ YOE3+ years building low-latency, high-throughput inference serving systems; knowledge of quantization, batching, speculative decoding, KV cache; experience with vLLM/TensorRT-LLM/Triton/SGLang; multi-GPU model parallelism; C++/CUDA/Python; autoscaling and GPU cost optimization.
ByteDance: Developing AI-driven content platforms and mobile applications.
PhD in a technical field; deep expertise in distributed storage and KV caching for LLMs, transformer internals, low-latency serving, GPU environments, and systems programming in C++/Rust/Go/CUDA.
Machine Learning Engineer, Inference & Serving (Speech LLM) - San Francisco
San Francisco, California, United States
$180k-$270k/yrHybridFull Time
Plaud: Develops AI-powered voice recorders and automated transcription software.
Experience building and deploying high-throughput, ultra-low-latency inference for LLMs or speech models; optimize latency/throughput; manage KV cache; understand GPU memory hierarchies; collaborate across ML and backend teams.
Discernis: AI-native document intelligence software for legal investigations and discovery.
Experience building production LLM/ML features end-to-end, hands-on with agentic orchestration, self-hosted model serving, strong Python and evaluation design skills.
ZeroDrift: AI-native compliance enforcement infrastructure for enterprise communication.
8+ YOE8+ years software engineering experience with 4+ years building ML/LLM infrastructure, hands-on PyTorch and LLM stack, production model serving, strong Python and cloud skills.
Vantaca: AI software for community association and HOA management.
8+ YOE8+ years in infrastructure/DevOps/SRE; strong cloud expertise; experience with CI/CD, PostgreSQL, Redis, APM, model serving, vector databases, GPU optimization, and LLM deployment.
PostgreSQL, Redis, APM, CI/CD, vector databases, model serving frameworks, LLM
Saviynt: Provides AI-powered identity governance and cloud security platforms.
ML platform or MLOps engineer with production Ray experience; LLM serving, distributed training, Python and PyTorch; MLflow/Flyte; Bachelor's degree in CS/Engineering.
Ray Train, Ray Serve, Ray Core, Ray Data, vLLM, SGLang, NVIDIA Triton, TorchTrainer, DDP, NCCL, PPO, RLlib, Flyte, MLflow, Qdrant, Pgvector, PyTorch, Python
NTT DATA: Global provider of IT and business consulting services.
10+ YOE10+ years in AI/ML, distributed systems, enterprise architecture or platform engineering; deep LLM, RAG, embeddings, model-serving and AWS experience; Terraform/IaC and DevOps pipeline expertise; ability to influence senior stakeholders.
Clera: AI talent agent matching professionals with high-growth startup roles
4+ YOE4+ years applied ML engineering in production, experience with LLMs/fine-tuning/RAG or large-scale recommender systems, strong Python and PyTorch or JAX skills, distributed training/GPU/inference serving experience, and mentoring ability.
Fireworks AI: Provides high-performance generative AI model inference and deployment infrastructure.
5+ YOE5+ years in customer-facing technical engineering roles, strong Python and Kubernetes skills, experience with LLM inference, model serving and fine-tuning, cloud GPU deployment across major clouds, and exceptional communication.
Python, Kubernetes, vLLM, SGLang, TensorRT-LLM, AWS, Microsoft Azure, GCP, Azure AI Foundry, AWS Bedrock, SageMaker, GCP Vertex
Nace AI: An enterprise AI product and research building long-running AI agents and specialized models for enterprise deployments.
5+ YOE5+ years MLOps or ML infrastructure experience, BS in CS or related, strong Python, Kubernetes, Docker, Terraform, GPU cluster and LLM serving experience, proficiency with orchestration and observability tools.
MassMutual: Sells life insurance, retirement plans, and investment services.
2+ YOE2+ years platform/infrastructure/SRE experience; cloud-native experience (AWS, GCP, or Azure); delivered platform features to production; experience with IaC and GitOps (Terraform or Pulumi, ArgoCD); familiarity with Kubernetes and LLM serving frameworks.
CapgeminiEuronext Paris: CAP: Provides global IT consulting and digital transformation services.
Experience leading AI platform or infrastructure teams, hands-on LLM inference/serving, cloud (AWS/Azure/GCP) with Kubernetes and IaC, regulated-industry delivery, research fluency, and platform-as-product mindset.
Lang Smith, Lang Graph Platform, model gateway, vector search, PostgreSQL, pg vector, ClickHouse, S3-compatible object storage, Neo4j Enterprise, Open Telemetry, Grafana, Kubernetes, Helm, Argo CD, AWS, Azure, GCP
Sprinter Health: Mobile provider of in-home diagnostic and preventive healthcare services.
8+ YOE8+ years building production ML systems and infrastructure; experience with training/serving pipelines, feature pipelines, monitoring, deployment, cloud, containers, CI/CD, and model governance.
DigitalOceanNew York Stock Exchange: DOCN: Simplifies cloud infrastructure for developers, startups, and SMBs.
Design and deliver high-scale, resilient data-plane services for inference; experience with LLM serving, distributed systems, and AI hardware; expert in Go or Python and gRPC; familiarity with inference frameworks and observability.
Chicago or Hong Kong or London or New York City or Singapore
$200k-$300k/yrOnsiteFull Time
DV Trading: Proprietary trading firm providing liquidity to global financial markets.
5+ YOE5+ years software engineering with strong Python; production fine-tuning/distillation of open-weight models; on-prem LLM serving, GPU infrastructure and Kubernetes experience; model evaluation and tooling.
Python, Llama, Qwen, Mistral, vLLM, TGI, Triton, Kubernetes, Hugging Face
Alt: Platform for trading, vaulting, and financing alternative assets.
7+ YOE7+ years engineering experience with 5+ years shipping production ML; strong Python and SQL; experience with MLflow, AWS, LLMs, production model serving, CI/CD, and feature engineering.