2,097 llm engineer jobs at 777 companies in California

2mo
Save
Mark Applied
Hide
LLM Inference Engineer
Los Altos, California, United States
OnsiteFull Time
Majestic Labs
Majestic Labs: Developing memory-first AI server platforms for data centers.
3+ YOE3+ years building or operating production LLM inference systems; strong Python and C++; experience with vLLM/SGLang/TensorRT-LLM/Fireworks; distributed inference and performance profiling skills.
vLLM, SGLang, TensorRT-LLM, Fireworks, Python, C++, collective communication library (CCL)
1mo
Save
Mark Applied
Hide
Principal LLM Inference Engineer
Santa Clara, California, United States
$195k-$285k/yr HybridFull Time
d-Matrix: Develops high-performance semiconductor chips for generative AI inference.
10+ YOEBachelor's in CS/EE (or equivalent) with 10+ years experience (Master/PhD with 6+ years preferred); strong Python and C/C++; experience optimizing LLM inference, quantization, batching, GPU kernel programming and contributor-level work on inference frameworks.
Python, C, C++, vLLM, SGLang, TensorRT-LLM, ONNX Runtime, CUDA, Triton, JAX
1w
Save
Mark Applied
Hide
Applied LLM Systems Engineer
Costa Mesa or Santa Ana
$112k-$149k/yr OnsiteFull Time
Anduril Industries
Anduril Industries: Defense technology building autonomous military hardware and software.
5+ YOERequires 5+ years of software engineering, production AI delivery, Python proficiency, LLM optimization, evaluation and observability experience, and eligibility for Secret or higher US security clearance.
Python, Lattice OS, S1000D, DITA, MIL-STD-40051, CI/CD
3mo
Save
Mark Applied
Hide
Distributed LLM Inference Engineer
San Francisco or Palo Alto
$170k-$247k/yr HybridFull Time
Anyscale
Anyscale: Cloud platform for scaling distributed machine learning applications.
Familiarity with running ML inference at large scale with high throughput and low latency; experience with PyTorch; solid understanding of distributed systems.
PyTorch, Ray, vLLM, TensorRT-LLM
1mo
Save
Mark Applied
Hide
Senior/Staff LLM Application Engineer - Data Application
San Jose, California, United States
$213k-$450k/yr OnsiteFull Time
TikTok
TikTok: Global short-form video hosting and social media platform.
Experience with data products and LLM application development, strong coding skills in Python, knowledge of prompt engineering, retrieval and benchmarking, and ability to analyze user feedback.
Python
3w
Save
Mark Applied
Hide
Software Engineer, AI Agent & LLM
Mountain View, California, United States
$155k-$185k/yr HybridFull Time
Otter.ai
Otter.ai: AI-powered meeting transcription and automated note-taking platform.
3+ YOE3+ years building AI-agent or ML systems, strong backend/distributed-systems engineering, experience shipping LLM-powered products, evaluation of nondeterministic systems, and ability to diagnose model and system failures.
3mo
Save
Mark Applied
Hide
LLM Solutions Architect
California or United States or Canada
$85k-$120k/yr HybridFull Time
Xsolla
Xsolla: Provides payment and monetization tools for game developers.
5+ YOE5+ years engineering; 2+ years deploying LLM systems; experience with major LLM APIs; Python; RAG pipelines; product ownership; strong communication.
Python, OpenAI, Anthropic, Google Gemini, LangChain, LlamaIndex, Vector databases, AutoGen, CrewAI
2mo
Save
Mark Applied
Hide
Machine Learning Engineer, LLM Post-Training
Mountain View, California, United States
$150k-$230k/yr OnsiteFull Time
NewsBreak
NewsBreak: Local news aggregation platform powered by artificial intelligence.
Hands-on LLM post-training (CPT, SFT, RL) with demonstrated RL experience; strong ML data engineering; experience training LLMs on mid-to-large GPU clusters; PyTorch and related frameworks familiarity; strong communication.
PyTorch, Hugging Face TRL, Hugging Face Accelerate, DeepSpeed, FSDP, vLLM
1w
Save
Mark Applied
Hide
LLM Backend Engineer Graduate (Applied Machine Learning) - 2027 Start
San Jose, California, United States
OnsiteFull Time
ByteDance
ByteDance: Developing AI-driven content platforms and mobile applications.
Bachelor's or master's degree in computer science or related field; proficiency in Golang, Java, C++, or Python; software development experience; and knowledge of databases, networks, operating systems, and distributed systems.
Golang, Java, C++, Python, Kubernetes, Docker, Istio, Envoy, Service Mesh, Function Calling, MCP
3w
Save
Mark Applied
Hide
Senior Machine Learning Engineer, LLM Inference Optimization
Palo Alto or California
$195k-$262k/yr OnsiteFull Time
Nebius
NebiusNasdaq: NBIS: Builds cloud infrastructure and software for artificial intelligence development.
Expert Python and PyTorch skills, hands-on LLM/VLM inference deployment and optimization, knowledge of modern inference stacks, quantitative reasoning about latency/throughput/cost, and strong communication.
Python, PyTorch, vLLM, SGLang, TensorRT-LLM, Triton Inference Server, NVIDIA Dynamo, Ray Serve, KServe, CUDA, FlashInfer, LMCache, Ray
3mo
Save
Mark Applied
Hide
Staff Machine Learning Engineer - Agentic Models, LLM, RAG, GenAI
Santa Clara, California, United States
$232k-$310k/yr HybridFull Time
Eightfold
Eightfold: Global AI-native talent intelligence platform provider.
6+ YOELead AI/ML engineer with 6+ years in ML, Gen AI, LLMs; strong Python, TensorFlow/PyTorch; AWS, Docker, Kubernetes; expert in agentic AI and distributed systems.
Python, TensorFlow, PyTorch, AWS, Docker, Kubernetes
3mo
Save
Mark Applied
Hide
Lead Machine Learning Engineer - Agentic Models, LLM, RAG, GenAI
Santa Clara, California, United States
$193k-$258k/yr HybridFull Time
Eightfold.ai
Eightfold.ai: AI-native platform for talent management and workforce optimization.
5+ YOESenior ML engineer with expertise in AI agents, LLMs, distributed systems; 5-7+ years of experience; strong Python and ML frameworks; AWS; Docker/Kubernetes; RAG/GenAI experience.
Python, TensorFlow, PyTorch, AWS, Docker, Kubernetes, Kafka, AWS SQS, LangGraph, CrewAI, AutoGen, Pinecone, pgvector, vLLM, TensorRT-LLM
2w
Save
Mark Applied
Hide
Senior Performance Co-Design Engineer, LLM Serving
Sunnyvale, California, United States
$174k-$252k/yr OnsiteFull Time
Google
GoogleNASDAQ: GOOGL: Provides online search, advertising, cloud computing, and consumer electronics.
5+ YOEBachelor's in CS/EE/CE or equivalent,5+ years in performance modeling/engineering or architecture,proficiency with C++ or Python,experience with ML serving and hardware/software co-design preferred.
C++, Python, TPU, Vertex AI
1mo
Save
Mark Applied
Hide
Senior Lead AI Engineer (FM Hosting, LLM Inference)
New York City or McLean or San Jose or Cambridge
$230k-$286k/yr OnsiteFull Time
Capital One
Capital OneNYSE: COF: Provides credit card, banking, and auto loan services.
6+ YOEBachelor’s plus 6 years or master’s plus 4 years in AI/ML development; 6 years programming in Python, Go, Scala, or Java; cloud AI deployment and engineering leadership preferred.
AWS Ultraclusters, Hugging Face, VectorDBs, Nemo Guardrails, PyTorch, Python, Go, Scala, Java, AWS, Google Cloud, Azure, C++, C#, Golang
1d
Save
Mark Applied
Hide
Senior Research Engineer, LLM Training & Post-Training
New York City or San Francisco or Seattle or London
$165k-$310k/yr HybridFull Time
Lightning AI
Lightning AI: Unified platform to build, train, and deploy AI models.
Requires significant PyTorch LLM training experience, distributed multi-GPU systems expertise, Python software engineering, experiment design, and a master's degree, PhD, or equivalent experience in a related field.
PyTorch, Python, DeepSpeed, FSDP, Megatron-LM, Hugging Face Transformers, TRL, PEFT, Lightning Fabric, CUDA, Triton, vLLM, SGLang, TensorRT, DPO, PPO, GRPO, SFT, RLHF
16h
Save
Mark Applied
Hide
Staff Machine Learning Engineer - LLM Quantization & Deployment
Santa Clara or Mountain View
$215k-$364k/yr OnsiteFull Time
XPeng
XPengNew York Stock Exchange: XPEV: Designs and manufactures smart electric vehicles and autonomous technology.
3+ YOEMaster's in CS, CE, or EE with 3–5 years' industry experience; expertise in Transformer architectures, LLM inference, model quantization, PyTorch, inference stacks, Python, and software engineering.
Python, PyTorch, TensorRT-LLM, vLLM, SGLang, llama.cpp, ONNX Runtime, TVM, MLIR, AWQ, GPTQ, SmoothQuant, INT8, FP4
1w
Save
Mark Applied
Hide
Founding AI Engineer
San Francisco, California, United States
$225k-$255k/yr OnsiteFull Time
Clera
Clera: AI talent agent matching professionals with high-growth startup roles
8+ YOE8+ years engineering experience with production LLM systems, building evals and observability, experience with LLM agents and data-residency/SOC2/GDPR constraints, strong communication and product-engineer instincts.
LLM, MCP, Anthropic, Google, LangChain, LlamaIndex, Braintrust, OpenRouter
1mo
Save
Mark Applied
Hide
Engineering Manager, LLM Performance
Santa Clara, California, United States
$224k-$431k/yr HybridFull Time
NVIDIA
NVIDIANASDAQ: NVDA: Designs GPU-accelerated computing and artificial intelligence hardware.
7+ YOE3+ MgmtMS/PhD or equivalent experience in CS/CE/AI, 7+ years software engineering experience including 3+ years technical leadership; strong C++ or Python; expertise in LLM/VLM/inference and production-quality software.
TensorRT LLM, vLLM, SGLang, Dynamo, C++, Python, CUDA
3mo
Save
Mark Applied
Hide
Machine Learning Engineer, Model Evaluations (Speech LLM) - San Francisco
San Francisco, California, United States
$180k-$270k/yr HybridFull Time
Plaud
Plaud: Develops AI-powered voice recorders and automated transcription software.
Python software engineering; building distributed systems, data pipelines, and evaluation harnesses at scale; partner with ML researchers to define benchmarks; build dashboards and monitor model health; debug mid-training anomalies; communicate results clearly.
Python, Distributed systems, Data pipelines, Evaluation harnesses, Dashboards, Weighs & Biases, MLflow
1mo
Save
Mark Applied
Hide
Engineering Manager, LLM Performance
Santa Clara, California, United States
$224k-$431k/yr HybridFull Time
NVIDIA
NVIDIANASDAQ: NVDA: Designs graphics processing units and artificial intelligence hardware.
7+ YOE3+ MgmtMS/PhD or equivalent in CS/CE/AI, 7+ years software engineering experience including 3+ years technical leadership, strong C++ or Python skills, expertise in LLM/VLM/inference and delivering production-quality software.
TensorRT LLM, TensorRT-LLM, vLLM, SGLang, Dynamo, C++, Python, CUDA