1,214 llm engineer jobs at 554 companies in Ross, CA
1mo
Save
Mark Applied
Hide
1mo
LLM Inference Engineer
Los Altos, California, United States
OnsiteFull Time
Majestic Labs: Developing memory-first AI server platforms for data centers.
3+ YOE3+ years building or operating production LLM inference systems; strong Python and C++; experience with vLLM/SGLang/TensorRT-LLM/Fireworks; distributed inference and performance profiling skills.
vLLM, SGLang, TensorRT-LLM, Fireworks, Python, C++, collective communication library (CCL)
Anyscale: Cloud platform for scaling distributed machine learning applications.
Familiarity with running ML inference at large scale with high throughput and low latency; experience with PyTorch; solid understanding of distributed systems.
Otter.ai: AI-powered meeting transcription and automated note-taking platform.
3+ YOE3+ years building AI-agent or ML systems, strong backend/distributed-systems engineering, experience shipping LLM-powered products, evaluation of nondeterministic systems, and ability to diagnose model and system failures.
NewsBreak: Local news aggregation platform powered by artificial intelligence.
Hands-on LLM post-training (CPT, SFT, RL) with demonstrated RL experience; strong ML data engineering; experience training LLMs on mid-to-large GPU clusters; PyTorch and related frameworks familiarity; strong communication.
PyTorch, Hugging Face TRL, Hugging Face Accelerate, DeepSpeed, FSDP, vLLM
NebiusNasdaq: NBIS: Builds cloud infrastructure and software for artificial intelligence development.
Expert Python and PyTorch skills, hands-on LLM/VLM inference deployment and optimization, knowledge of modern inference stacks, quantitative reasoning about latency/throughput/cost, and strong communication.
Python, PyTorch, vLLM, SGLang, TensorRT-LLM, Triton Inference Server, NVIDIA Dynamo, Ray Serve, KServe, CUDA, FlashInfer, LMCache, Ray
5+ YOEBachelor's in CS/EE/CE or equivalent,5+ years in performance modeling/engineering or architecture,proficiency with C++ or Python,experience with ML serving and hardware/software co-design preferred.
Clera: AI talent agent matching professionals with high-growth startup roles
8+ YOE8+ years engineering experience with production LLM systems, building evals and observability, experience with LLM agents and data-residency/SOC2/GDPR constraints, strong communication and product-engineer instincts.
Machine Learning Engineer, Model Evaluations (Speech LLM) - San Francisco
San Francisco, California, United States
$180k-$270k/yrHybridFull Time
Plaud: Develops AI-powered voice recorders and automated transcription software.
Python software engineering; building distributed systems, data pipelines, and evaluation harnesses at scale; partner with ML researchers to define benchmarks; build dashboards and monitor model health; debug mid-training anomalies; communicate results clearly.
Mira Mace: AI-powered healthcare advocacy and navigation for Medicare beneficiaries.
Staff-level backend/full-stack or ML engineering experience, production LLM/ML systems experience, architecture and build-vs-buy judgment, hands-on coding and code review, mentoring and technical leadership.
Baseten: Scalable infrastructure platform for deploying and serving AI models.
4+ YOE1+ MgmtLead a team of Forward Deployed Engineers; strong Python, ML inference, LLM experience; 4+ years software engineering; leadership experience; excellent communication.
Python, vLLM, TensorRT, Triton, Hugging Face, Ray Serve
Shakudo: Develops an operating system for enterprise AI applications.
8+ YOE8+ years engineering experience, 5+ years Kubernetes operation, proficiency in Rust, experience with production infrastructure (physical servers, GPU/DGX clusters), CI/CD, security hardening, observability, and LLM/AI infrastructure.
Beacon AI: Developing an AI-powered-pilot for safer flight operations.
Experience building LLM-powered features, RAG and tool-calling, production services in Python or TypeScript, vector search and embeddings, evals/metrics, and safety/compliance for a regulated domain.
San Francisco or New York City or Vancouver or Amsterdam or London or Paris or Singapore or Tokyo
$190k-$316k/yrHybridFull Time
AmplitudeNASDAQ: AMPL: Digital analytics platform for understanding and optimizing customer behavior.
3+ YOE3+ years software engineering experience with 2+ years building production LLM or applied AI systems; experience designing and scaling AI products, strong product sense, and collaboration skills.
Austin or Chicago or New York City or Salt Lake City or San Francisco or United States
$115k-$175k/yrRemoteFull Time
Gong: AI platform analyzing customer interactions to improve sales performance.
Background in software/data engineering or applied AI, experience shipping AI/LLM tools to production, systems thinking, API and workflow integration experience, strong independent delivery and stakeholder partnership skills.
AdobeNASDAQ: ADBE: Provides software for digital media creation and marketing analytics
10+ YOE10+ years at the intersection of design and engineering; hands-on experience building and shipping LLM-powered apps, internal tools, and high-fidelity prototypes; strong craft, collaboration, and product judgment.
Mondrio: AI-native platform for agentic pricing and revenue management.
8+ YOE8+ years engineering experience with production LLMs, built evals/observability, shipped LLM features, strong communication, product-engineer instincts, US work authorization required.
Databricks: A unified platform for data analytics and artificial intelligence.
12+ YOE12+ years software engineering experience, production LLM systems experience, human-in-the-loop design, CMS and web publishing pipeline knowledge, stakeholder collaboration, and mentoring experience.
San Francisco or Toronto or United States or Canada
$190k-$249k/yrRemoteFull Time
EvenUp: AI-powered document generation and analysis for personal injury law.
5+ YOE5+ years engineering experience, building production-quality software; ownership of projects end-to-end; experience with frontend systems for deploying LLM-based algorithms; strong communication, mentoring, and cross-team collaboration skills.
Model AI: Building high-performance infrastructure for agentic AI systems.
Strong Python and systems engineering skills, experience building LLM agents, evaluation harnesses, CI/CD and developer tooling, testing and benchmarking, and ability to turn research into production systems.