1,439 llm engineer jobs at 580 companies in Benicia, CA
1mo
Save
Mark Applied
Hide
1mo
LLM Inference Engineer
Los Altos, California, United States
OnsiteFull Time
Majestic Labs: Developing memory-first AI server platforms for data centers.
3+ YOE3+ years building or operating production LLM inference systems; strong Python and C++; experience with vLLM/SGLang/TensorRT-LLM/Fireworks; distributed inference and performance profiling skills.
vLLM, SGLang, TensorRT-LLM, Fireworks, Python, C++, collective communication library (CCL)
d-Matrix: Develops high-performance semiconductor chips for generative AI inference.
10+ YOEBachelor's in CS/EE (or equivalent) with 10+ years experience (Master/PhD with 6+ years preferred); strong Python and C/C++; experience optimizing LLM inference, quantization, batching, GPU kernel programming and contributor-level work on inference frameworks.
Anyscale: Cloud platform for scaling distributed machine learning applications.
Familiarity with running ML inference at large scale with high throughput and low latency; experience with PyTorch; solid understanding of distributed systems.
Otter.ai: AI-powered meeting transcription and automated note-taking platform.
3+ YOE3+ years building AI-agent or ML systems, strong backend/distributed-systems engineering, experience shipping LLM-powered products, evaluation of nondeterministic systems, and ability to diagnose model and system failures.
NewsBreak: Local news aggregation platform powered by artificial intelligence.
Hands-on LLM post-training (CPT, SFT, RL) with demonstrated RL experience; strong ML data engineering; experience training LLMs on mid-to-large GPU clusters; PyTorch and related frameworks familiarity; strong communication.
PyTorch, Hugging Face TRL, Hugging Face Accelerate, DeepSpeed, FSDP, vLLM
NebiusNasdaq: NBIS: Builds cloud infrastructure and software for artificial intelligence development.
Expert Python and PyTorch skills, hands-on LLM/VLM inference deployment and optimization, knowledge of modern inference stacks, quantitative reasoning about latency/throughput/cost, and strong communication.
Python, PyTorch, vLLM, SGLang, TensorRT-LLM, Triton Inference Server, NVIDIA Dynamo, Ray Serve, KServe, CUDA, FlashInfer, LMCache, Ray
Eightfold: Global AI-native talent intelligence platform provider.
6+ YOELead AI/ML engineer with 6+ years in ML, Gen AI, LLMs; strong Python, TensorFlow/PyTorch; AWS, Docker, Kubernetes; expert in agentic AI and distributed systems.
Lead Machine Learning Engineer - Agentic Models, LLM, RAG, GenAI
Santa Clara, California, United States
$193k-$258k/yrHybridFull Time
Eightfold.ai: AI-native platform for talent management and workforce optimization.
5+ YOESenior ML engineer with expertise in AI agents, LLMs, distributed systems; 5-7+ years of experience; strong Python and ML frameworks; AWS; Docker/Kubernetes; RAG/GenAI experience.
5+ YOEBachelor's in CS/EE/CE or equivalent,5+ years in performance modeling/engineering or architecture,proficiency with C++ or Python,experience with ML serving and hardware/software co-design preferred.
Senior Research Engineer, LLM Training & Post-Training
New York City or San Francisco or Seattle or London
$165k-$310k/yrHybridFull Time
Lightning AI: Unified platform to build, train, and deploy AI models.
Requires significant PyTorch LLM training experience, distributed multi-GPU systems expertise, Python software engineering, experiment design, and a master's degree, PhD, or equivalent experience in a related field.
Clera: AI talent agent matching professionals with high-growth startup roles
8+ YOE8+ years engineering experience with production LLM systems, building evals and observability, experience with LLM agents and data-residency/SOC2/GDPR constraints, strong communication and product-engineer instincts.
Machine Learning Engineer, Model Evaluations (Speech LLM) - San Francisco
San Francisco, California, United States
$180k-$270k/yrHybridFull Time
Plaud: Develops AI-powered voice recorders and automated transcription software.
Python software engineering; building distributed systems, data pipelines, and evaluation harnesses at scale; partner with ML researchers to define benchmarks; build dashboards and monitor model health; debug mid-training anomalies; communicate results clearly.
NVIDIANASDAQ: NVDA: Designs GPU-accelerated computing and artificial intelligence hardware.
7+ YOE3+ MgmtMS/PhD or equivalent experience in CS/CE/AI, 7+ years software engineering experience including 3+ years technical leadership; strong C++ or Python; expertise in LLM/VLM/inference and production-quality software.
TensorRT LLM, vLLM, SGLang, Dynamo, C++, Python, CUDA
NVIDIANASDAQ: NVDA: Designs graphics processing units and artificial intelligence hardware.
7+ YOE3+ MgmtMS/PhD or equivalent in CS/CE/AI, 7+ years software engineering experience including 3+ years technical leadership, strong C++ or Python skills, expertise in LLM/VLM/inference and delivering production-quality software.
TensorRT LLM, TensorRT-LLM, vLLM, SGLang, Dynamo, C++, Python, CUDA
Mira Mace: AI-powered healthcare advocacy and navigation for Medicare beneficiaries.
Staff-level backend/full-stack or ML engineering experience, production LLM/ML systems experience, architecture and build-vs-buy judgment, hands-on coding and code review, mentoring and technical leadership.
Baseten: Scalable infrastructure platform for deploying and serving AI models.
4+ YOE1+ MgmtLead a team of Forward Deployed Engineers; strong Python, ML inference, LLM experience; 4+ years software engineering; leadership experience; excellent communication.
Python, vLLM, TensorRT, Triton, Hugging Face, Ray Serve
McLean or Richmond or New York City or San Jose or Cambridge or Plano or San Francisco
$245k-$335k/yrOnsiteFull Time
Capital OneNYSE: COF: Provides credit card, banking, and auto loan services.
7+ YOEBachelor's degree; 7+ years in software engineering and solution architecture; 5+ years shipping cloud platforms; 3+ years with LLM systems, RAG, embeddings, prompt tooling, and model governance.
Shakudo: Develops an operating system for enterprise AI applications.
8+ YOE8+ years engineering experience, 5+ years Kubernetes operation, proficiency in Rust, experience with production infrastructure (physical servers, GPU/DGX clusters), CI/CD, security hardening, observability, and LLM/AI infrastructure.
Beacon AI: Developing an AI-powered-pilot for safer flight operations.
Experience building LLM-powered features, RAG and tool-calling, production services in Python or TypeScript, vector search and embeddings, evals/metrics, and safety/compliance for a regulated domain.
San Francisco or New York City or Vancouver or Amsterdam or London or Paris or Singapore or Tokyo
$190k-$316k/yrHybridFull Time
AmplitudeNASDAQ: AMPL: Digital analytics platform for understanding and optimizing customer behavior.
3+ YOE3+ years software engineering experience with 2+ years building production LLM or applied AI systems; experience designing and scaling AI products, strong product sense, and collaboration skills.