3,086 llm jobs at 1,036 companies in California

3mo
Save
Mark Applied
Hide
LLM Solutions Architect
California or United States or Canada
$85k-$120k/yr HybridFull Time
Xsolla
Xsolla: Provides payment and monetization tools for game developers.
5+ YOE5+ years engineering; 2+ years deploying LLM systems; experience with major LLM APIs; Python; RAG pipelines; product ownership; strong communication.
Python, OpenAI, Anthropic, Google Gemini, LangChain, LlamaIndex, Vector databases, AutoGen, CrewAI
2mo
Save
Mark Applied
Hide
LLM Inference Engineer
Los Altos, California, United States
OnsiteFull Time
Majestic Labs
Majestic Labs: Developing memory-first AI server platforms for data centers.
3+ YOE3+ years building or operating production LLM inference systems; strong Python and C++; experience with vLLM/SGLang/TensorRT-LLM/Fireworks; distributed inference and performance profiling skills.
vLLM, SGLang, TensorRT-LLM, Fireworks, Python, C++, collective communication library (CCL)
3mo
Save
Mark Applied
Hide
Multimodal LLM Researcher (MLLM)
Palo Alto, California, United States
$185k-$400k/yr HybridFull Time
Pika
Pika: AI-powered platform for generating and editing professional videos
5+ YOE5+ years in LLM/VLM/Audio LM, deep learning, and related fields; publications in top venues; real-time generative models experience; strong Python and ML framework skills; dataset curation experience.
Python, PyTorch, TensorFlow
1mo
Save
Mark Applied
Hide
Principal LLM Inference Engineer
Santa Clara, California, United States
$195k-$285k/yr HybridFull Time
d-Matrix: Develops high-performance semiconductor chips for generative AI inference.
10+ YOEBachelor's in CS/EE (or equivalent) with 10+ years experience (Master/PhD with 6+ years preferred); strong Python and C/C++; experience optimizing LLM inference, quantization, batching, GPU kernel programming and contributor-level work on inference frameworks.
Python, C, C++, vLLM, SGLang, TensorRT-LLM, ONNX Runtime, CUDA, Triton, JAX
1mo
Save
Mark Applied
Hide
Senior/Staff LLM Application Engineer - Data Application
San Jose, California, United States
$213k-$450k/yr OnsiteFull Time
TikTok
TikTok: Global short-form video hosting and social media platform.
Experience with data products and LLM application development, strong coding skills in Python, knowledge of prompt engineering, retrieval and benchmarking, and ability to analyze user feedback.
Python
3w
Save
Mark Applied
Hide
Software Engineer, AI Agent & LLM
Mountain View, California, United States
$155k-$185k/yr HybridFull Time
Otter.ai
Otter.ai: AI-powered meeting transcription and automated note-taking platform.
3+ YOE3+ years building AI-agent or ML systems, strong backend/distributed-systems engineering, experience shipping LLM-powered products, evaluation of nondeterministic systems, and ability to diagnose model and system failures.
3mo
Save
Mark Applied
Hide
Distributed LLM Inference Engineer
San Francisco or Palo Alto
$170k-$247k/yr HybridFull Time
Anyscale
Anyscale: Cloud platform for scaling distributed machine learning applications.
Familiarity with running ML inference at large scale with high throughput and low latency; experience with PyTorch; solid understanding of distributed systems.
PyTorch, Ray, vLLM, TensorRT-LLM
1w
Save
Mark Applied
Hide
LLM Backend Engineer Graduate (Applied Machine Learning) - 2027 Start
San Jose, California, United States
OnsiteFull Time
ByteDance
ByteDance: Developing AI-driven content platforms and mobile applications.
Bachelor's or master's degree in computer science or related field; proficiency in Golang, Java, C++, or Python; software development experience; and knowledge of databases, networks, operating systems, and distributed systems.
Golang, Java, C++, Python, Kubernetes, Docker, Istio, Envoy, Service Mesh, Function Calling, MCP
1mo
Save
Mark Applied
Hide
Engineering Manager, LLM Performance
Santa Clara, California, United States
$224k-$431k/yr HybridFull Time
NVIDIA
NVIDIANASDAQ: NVDA: Designs GPU-accelerated computing and artificial intelligence hardware.
7+ YOE3+ MgmtMS/PhD or equivalent experience in CS/CE/AI, 7+ years software engineering experience including 3+ years technical leadership; strong C++ or Python; expertise in LLM/VLM/inference and production-quality software.
TensorRT LLM, vLLM, SGLang, Dynamo, C++, Python, CUDA
1mo
Save
Mark Applied
Hide
Engineering Manager, LLM Performance
Santa Clara, California, United States
$224k-$431k/yr HybridFull Time
NVIDIA
NVIDIANASDAQ: NVDA: Designs graphics processing units and artificial intelligence hardware.
7+ YOE3+ MgmtMS/PhD or equivalent in CS/CE/AI, 7+ years software engineering experience including 3+ years technical leadership, strong C++ or Python skills, expertise in LLM/VLM/inference and delivering production-quality software.
TensorRT LLM, TensorRT-LLM, vLLM, SGLang, Dynamo, C++, Python, CUDA
1mo
Save
Mark Applied
Hide
AI Researcher, On-Device LLM Efficiency
San Diego, California, United States
$139k-$208k/yr OnsiteFull Time
Qualcomm
QualcommNASDAQ: QCOM: Designs and manufactures semiconductors and wireless telecommunications products.
4+ YOEMaster's in CS/EE or related, 4+ years AI research experience (LLM/Transformers), strong deep learning background, Python and PyTorch skills, experience with LLM inference or on-device deployment.
Python, PyTorch, Transformers
2mo
Save
Mark Applied
Hide
Machine Learning Engineer, LLM Post-Training
Mountain View, California, United States
$150k-$230k/yr OnsiteFull Time
NewsBreak
NewsBreak: Local news aggregation platform powered by artificial intelligence.
Hands-on LLM post-training (CPT, SFT, RL) with demonstrated RL experience; strong ML data engineering; experience training LLMs on mid-to-large GPU clusters; PyTorch and related frameworks familiarity; strong communication.
PyTorch, Hugging Face TRL, Hugging Face Accelerate, DeepSpeed, FSDP, vLLM
3mo
Save
Mark Applied
Hide
Research Scientist Intern, Monetization Generative AI - LLM (PhD)
Bellevue or Menlo Park or Seattle or New York
$8k-$12k/mo HybridInternship
Meta
MetaNASDAQ: META: Develops social networking platforms and virtual reality technologies.
Ph.D. in CS/AI/NLP/Speech/CV or related field; work authorization; experience in Python/C/C++/Java; experience with PyTorch or TensorFlow; ML/NLP system building.
Python, C++, C, Java, PyTorch, TensorFlow
3mo
Save
Mark Applied
Hide
Engineering Manager - Forward Deployed Engineering (LLM)
San Francisco or New York
$260k-$380k/yr HybridFull Time
Baseten
Baseten: Scalable infrastructure platform for deploying and serving AI models.
4+ YOE1+ MgmtLead a team of Forward Deployed Engineers; strong Python, ML inference, LLM experience; 4+ years software engineering; leadership experience; excellent communication.
Python, vLLM, TensorRT, Triton, Hugging Face, Ray Serve
1w
Save
Mark Applied
Hide
Senior Software Development Engineer in Test — LLM Evaluation & Automation, T3E
San Diego, California, United States
OnsiteFull Time
Apple
AppleNASDAQ: AAPL: Designs and sells consumer electronics, software, and online services.
Senior hands-on individual contributor leading automated LLM model evaluation, building regression-detection infrastructure, and partnering with modeling, framework, and infrastructure teams.
LLM
1mo
Save
Mark Applied
Hide
AI/LLM Product Director - Executive Director
Palo Alto or New York
$181k-$285k/yr OnsiteFull Time
JPMorgan Chase
JPMorgan ChaseNYSE: JPM: Global financial services firm providing banking and investment solutions.
8+ YOE8+ years building and launching AI/LLM products; knowledge of ElasticSearch, LangChain/LangGraph/OpenLLM, SFT, RLHF, RAG, Agents, MLOps, cloud (AWS); strong collaboration and product leadership skills.
ElasticSearch, LangChain, LangGraph, OpenLLM, AWS, MLOps
3w
Save
Mark Applied
Hide
Software Development Manager, LLM Inference Model Enablement, Neuron SDK
Cupertino, California, United States
$213k-$288k/yr OnsiteFull Time
Amazon
AmazonNASDAQ: AMZN: Global online retail and cloud computing technology provider.
7+ YOE3+ MgmtManage engineering team to onboard and optimize LLMs for inference on Trainium; strong background in LLM architectures, model performance optimization, and inference techniques; experience with PyTorch and Neuron stack.
PyTorch, AWS Neuron, Neuron compiler, Neuron runtime
1d
Save
Mark Applied
Hide
Staff Machine Learning Engineer - LLM Quantization & Deployment
Santa Clara or Mountain View
$215k-$364k/yr OnsiteFull Time
XPeng
XPengNew York Stock Exchange: XPEV: Designs and manufactures smart electric vehicles and autonomous technology.
3+ YOEMaster's in CS, CE, or EE with 3–5 years' industry experience; expertise in Transformer architectures, LLM inference, model quantization, PyTorch, inference stacks, Python, and software engineering.
Python, PyTorch, TensorRT-LLM, vLLM, SGLang, llama.cpp, ONNX Runtime, TVM, MLIR, AWQ, GPTQ, SmoothQuant, INT8, FP4
3w
Save
Mark Applied
Hide
Senior Machine Learning Engineer, LLM Inference Optimization
Palo Alto or California
$195k-$262k/yr OnsiteFull Time
Nebius
NebiusNasdaq: NBIS: Builds cloud infrastructure and software for artificial intelligence development.
Expert Python and PyTorch skills, hands-on LLM/VLM inference deployment and optimization, knowledge of modern inference stacks, quantitative reasoning about latency/throughput/cost, and strong communication.
Python, PyTorch, vLLM, SGLang, TensorRT-LLM, Triton Inference Server, NVIDIA Dynamo, Ray Serve, KServe, CUDA, FlashInfer, LMCache, Ray
2d
Save
Mark Applied
Hide
Senior Research Engineer, LLM Training & Post-Training
New York City or San Francisco or Seattle or London
$165k-$310k/yr HybridFull Time
Lightning AI
Lightning AI: Unified platform to build, train, and deploy AI models.
Requires significant PyTorch LLM training experience, distributed multi-GPU systems expertise, Python software engineering, experiment design, and a master's degree, PhD, or equivalent experience in a related field.
PyTorch, Python, DeepSpeed, FSDP, Megatron-LM, Hugging Face Transformers, TRL, PEFT, Lightning Fabric, CUDA, Triton, vLLM, SGLang, TensorRT, DPO, PPO, GRPO, SFT, RLHF