1,634 llm engineer jobs at 628 companies in Dublin, CA

PromotedHiringCafe
Founding Machine Learning / AI Search Engineer
Cupertino, CA, US
$160k-$310k/yr On-SiteFull Time
HiringCafe
HiringCafe: Building a 100× better job search engine to take on Indeed and LinkedIn.
Build the ML and AI search behind HiringCafe — ranking, recommenders, retrieval, and LLM agents that surface jobs people would never find on their own.
Python, PyTorch, Elasticsearch, LLMs
1mo
Save
Mark Applied
Hide
LLM Inference Engineer
Los Altos, California, United States
OnsiteFull Time
Majestic Labs
Majestic Labs: Developing memory-first AI server platforms for data centers.
3+ YOE3+ years building or operating production LLM inference systems; strong Python and C++; experience with vLLM/SGLang/TensorRT-LLM/Fireworks; distributed inference and performance profiling skills.
vLLM, SGLang, TensorRT-LLM, Fireworks, Python, C++, collective communication library (CCL)
2w
Save
Mark Applied
Hide
LLM Inference Engineer
San Francisco or United States
RemoteFull Time
NEAR AI
NEAR AI: Building private, verifiable infrastructure for autonomous AI agents.
Expert in LLM inference and serving systems, optimizing throughput/latency/cost for open-source LLMs, deep GPU architecture knowledge, and experience with PyTorch, Triton, CUDA and inference engines like vLLM/SGLang/TensorRT.
SGLang, vLLM, TensorRT, PyTorch, Triton, CuTe, CUDA
3w
Save
Mark Applied
Hide
Principal LLM Inference Engineer
Santa Clara, California, United States
$195k-$285k/yr HybridFull Time
d-Matrix: Develops high-performance semiconductor chips for generative AI inference.
10+ YOEBachelor's in CS/EE (or equivalent) with 10+ years experience (Master/PhD with 6+ years preferred); strong Python and C/C++; experience optimizing LLM inference, quantization, batching, GPU kernel programming and contributor-level work on inference frameworks.
Python, C, C++, vLLM, SGLang, TensorRT-LLM, ONNX Runtime, CUDA, Triton, JAX
2mo
Save
Mark Applied
Hide
Distributed LLM Inference Engineer
San Francisco or Palo Alto
$170k-$247k/yr HybridFull Time
Anyscale
Anyscale: Cloud platform for scaling distributed machine learning applications.
Familiarity with running ML inference at large scale with high throughput and low latency; experience with PyTorch; solid understanding of distributed systems.
PyTorch, Ray, vLLM, TensorRT-LLM
3mo
Save
Mark Applied
Hide
Sr. LLM Application Engineer (AI Agents)
Palo Alto, California, United States
$145k-$273k/yr OnsiteFull Time
Tencent
TencentHong Kong Stock Exchange: 0700: Developing digital services and entertainment for a global audience.
Bachelor's degree in CS/AI/Software Engineering; strong experience in LLM application development, AI Agents, RAG optimization; knowledge of MAS, tool-use protocols; production AI agent systems.
2w
Save
Mark Applied
Hide
Senior/Staff LLM Application Engineer - Data Application
San Jose, California, United States
$213k-$450k/yr OnsiteFull Time
TikTok
TikTok: Global short-form video hosting and social media platform.
Experience with data products and LLM application development, strong coding skills in Python, knowledge of prompt engineering, retrieval and benchmarking, and ability to analyze user feedback.
Python
2mo
Save
Mark Applied
Hide
Principal High-Performance LLM Training Engineer
Santa Clara, California, United States
$272k-$431k/yr OnsiteFull Time
NVIDIA
NVIDIANASDAQ: NVDA: Designs graphics processing units and artificial intelligence hardware.
12+ YOEMS/PhD or equivalent with 12+ years experience; principal-level impact in large-scale AI training, GPU and distributed systems performance; experience with transformer/LLM workloads, distributed training techniques, profiling and benchmarking; strong communication and leadership.
PyTorch, JAX, NeMo, NeMo RL, CUDA
1d
Save
Mark Applied
Hide
Software Engineer, AI Agent & LLM
Mountain View, California, United States
$155k-$185k/yr HybridFull Time
Otter.ai
Otter.ai: AI-powered meeting transcription and automated note-taking platform.
3+ YOE3+ years building AI-agent or ML systems, strong backend/distributed-systems engineering, experience shipping LLM-powered products, evaluation of nondeterministic systems, and ability to diagnose model and system failures.
2mo
Save
Mark Applied
Hide
Principal High-Performance LLM Training Engineer
Santa Clara, California, United States
$272k-$431k/yr OnsiteFull Time
NVIDIA
NVIDIANASDAQ: NVDA: Designs GPU-accelerated computing and artificial intelligence hardware.
12+ YOEMS or PhD in CS/EE/CE, with 12+ years of relevant work; strong impact in large-scale AI training, GPU performance, distributed systems, or similar.
PyTorch, JAX, NeMo, NeMo RL, CUDA, Profiling tools
1mo
Save
Mark Applied
Hide
Machine Learning Engineer, LLM Post-Training
Mountain View, California, United States
$150k-$230k/yr OnsiteFull Time
NewsBreak
NewsBreak: Local news aggregation platform powered by artificial intelligence.
Hands-on LLM post-training (CPT, SFT, RL) with demonstrated RL experience; strong ML data engineering; experience training LLMs on mid-to-large GPU clusters; PyTorch and related frameworks familiarity; strong communication.
PyTorch, Hugging Face TRL, Hugging Face Accelerate, DeepSpeed, FSDP, vLLM
1d
Save
Mark Applied
Hide
Senior Machine Learning Engineer, LLM Inference Optimization
Palo Alto or California
$195k-$262k/yr OnsiteFull Time
Nebius
NebiusNasdaq: NBIS: Builds cloud infrastructure and software for artificial intelligence development.
Expert Python and PyTorch skills, hands-on LLM/VLM inference deployment and optimization, knowledge of modern inference stacks, quantitative reasoning about latency/throughput/cost, and strong communication.
Python, PyTorch, vLLM, SGLang, TensorRT-LLM, Triton Inference Server, NVIDIA Dynamo, Ray Serve, KServe, CUDA, FlashInfer, LMCache, Ray
2w
Save
Mark Applied
Hide
LLM AIOps Development Engineer - Data Center Networking
San Jose, California, United States
OnsiteFull Time
ByteDance
ByteDance: Developing AI-driven content platforms and mobile applications.
Deep data-center networking and Linux networking knowledge; proficiency in Golang or Python; experience with telemetry, observability, big-data pipelines, and LLM/AIOps approaches; familiarity with protocols like EVPN/VXLAN and BGP/OSPF.
gNMI, Netconf, IPFIX, NetFlow, SNMP, Golang, Python, Docker, Kubernetes, CI/CD, Kafka, Flink, ClickHouse, TSDB, Prometheus, OpenTelemetry, Neo4j, SONiC, P4, eBPF, DPDK, RDMA, RoCE
2mo
Save
Mark Applied
Hide
Staff Machine Learning Engineer - Agentic Models, LLM, RAG, GenAI
Santa Clara, California, United States
$232k-$310k/yr HybridFull Time
Eightfold
Eightfold: Global AI-native talent intelligence platform provider.
6+ YOELead AI/ML engineer with 6+ years in ML, Gen AI, LLMs; strong Python, TensorFlow/PyTorch; AWS, Docker, Kubernetes; expert in agentic AI and distributed systems.
Python, TensorFlow, PyTorch, AWS, Docker, Kubernetes
2mo
Save
Mark Applied
Hide
Lead Machine Learning Engineer - Agentic Models, LLM, RAG, GenAI
Santa Clara, California, United States
$193k-$258k/yr HybridFull Time
Eightfold.ai
Eightfold.ai: AI-native platform for talent management and workforce optimization.
5+ YOESenior ML engineer with expertise in AI agents, LLMs, distributed systems; 5-7+ years of experience; strong Python and ML frameworks; AWS; Docker/Kubernetes; RAG/GenAI experience.
Python, TensorFlow, PyTorch, AWS, Docker, Kubernetes, Kafka, AWS SQS, LangGraph, CrewAI, AutoGen, Pinecone, pgvector, vLLM, TensorRT-LLM
2mo
Save
Mark Applied
Hide
Machine Learning Engineer, LLM Evals & Observability
San Francisco or Mountain View
$200k-$300k/yr HybridFull Time
Glean
Glean: AI platform for enterprise search and automated workplace agents
2+ YOE2+ years software engineering; Go and Python; distributed data pipelines; LLM evaluation, RLHF, NLP; strong backend; analytics mindset
Go, Python, LLM evaluation, NLP, data pipelines
2mo
Save
Mark Applied
Hide
Software Engineering Manager, LLM Training
Mountain View, California, United States
$170k-$277k/yr HybridFull Time
LinkedInNASDAQ: MSFT: Professional social network for career development and job recruitment.
1+ Mgmt5+ years software engineering; 1+ year management; experience with LLMs and distributed systems; strong leadership and strategic planning
PyTorch, CUDA, Megatron, Hugging Face, vLLM, Liger, Ray, SGLang, VERL, NCCL, TorchDistributed
2mo
Save
Mark Applied
Hide
Machine Learning Engineer, Model Evaluations (Speech LLM) - San Francisco
San Francisco, California, United States
$180k-$270k/yr HybridFull Time
Plaud
Plaud: Develops AI-powered voice recorders and automated transcription software.
Python software engineering; building distributed systems, data pipelines, and evaluation harnesses at scale; partner with ML researchers to define benchmarks; build dashboards and monitor model health; debug mid-training anomalies; communicate results clearly.
Python, Distributed systems, Data pipelines, Evaluation harnesses, Dashboards, Weighs & Biases, MLflow
2mo
Save
Mark Applied
Hide
Engineering Manager - Forward Deployed Engineering (LLM)
San Francisco or New York
$260k-$380k/yr HybridFull Time
Baseten
Baseten: Scalable infrastructure platform for deploying and serving AI models.
4+ YOE1+ MgmtLead a team of Forward Deployed Engineers; strong Python, ML inference, LLM experience; 4+ years software engineering; leadership experience; excellent communication.
Python, vLLM, TensorRT, Triton, Hugging Face, Ray Serve
4w
Save
Mark Applied
Hide
Infrastructure Engineer
Menlo Park, California, United States
OnsiteFull Time
Shakudo
Shakudo: Develops an operating system for enterprise AI applications.
8+ YOE8+ years engineering experience, 5+ years Kubernetes operation, proficiency in Rust, experience with production infrastructure (physical servers, GPU/DGX clusters), CI/CD, security hardening, observability, and LLM/AI infrastructure.
Kubernetes, Rust, CI/CD, DGX, GPU, LLM, ETL
1w
Save
Mark Applied
Hide
Software Engineer, Artificial Intelligence/LLM (Multiple Seniority Levels)
San Carlos, California, United States
$135k-$260k/yr HybridFull Time
Beacon AI
Beacon AI: Developing an AI-powered-pilot for safer flight operations.
Experience shipping LLM-powered features, building RAG and tool-calling flows, production Python/TypeScript services, and working with embeddings, vector backends, and model providers.
LangChain, Python, TypeScript, AWS Bedrock, OpenAI, Anthropic, OpenSearch, OpenSearch Serverless, pgvector, Pinecone, Weaviate, S3, Aurora, DynamoDB, Triton, TensorRT-LLM, CI/CD, IaC