1,591 llm engineer jobs at 636 companies in Newark, CA
1mo
Save
Mark Applied
Hide
1mo
LLM Inference Engineer
Los Altos, California, United States
OnsiteFull Time
Majestic Labs: Developing memory-first AI server platforms for data centers.
3+ YOE3+ years building or operating production LLM inference systems; strong Python and C++; experience with vLLM/SGLang/TensorRT-LLM/Fireworks; distributed inference and performance profiling skills.
vLLM, SGLang, TensorRT-LLM, Fireworks, Python, C++, collective communication library (CCL)
NEAR AI: Building private, verifiable infrastructure for autonomous AI agents.
Expert in LLM inference and serving systems, optimizing throughput/latency/cost for open-source LLMs, deep GPU architecture knowledge, and experience with PyTorch, Triton, CUDA and inference engines like vLLM/SGLang/TensorRT.
SGLang, vLLM, TensorRT, PyTorch, Triton, CuTe, CUDA
d-Matrix: Develops high-performance semiconductor chips for generative AI inference.
10+ YOEBachelor's in CS/EE (or equivalent) with 10+ years experience (Master/PhD with 6+ years preferred); strong Python and C/C++; experience optimizing LLM inference, quantization, batching, GPU kernel programming and contributor-level work on inference frameworks.
Anyscale: Cloud platform for scaling distributed machine learning applications.
Familiarity with running ML inference at large scale with high throughput and low latency; experience with PyTorch; solid understanding of distributed systems.
TencentHong Kong Stock Exchange: 0700: Developing digital services and entertainment for a global audience.
Bachelor's degree in CS/AI/Software Engineering; strong experience in LLM application development, AI Agents, RAG optimization; knowledge of MAS, tool-use protocols; production AI agent systems.
Senior/Staff LLM Application Engineer - Data Application
San Jose, California, United States
$213k-$450k/yrOnsiteFull Time
TikTok: Global short-form video hosting and social media platform.
Experience with data products and LLM application development, strong coding skills in Python, knowledge of prompt engineering, retrieval and benchmarking, and ability to analyze user feedback.
Otter.ai: AI-powered meeting transcription and automated note-taking platform.
3+ YOE3+ years building AI-agent or ML systems, strong backend/distributed-systems engineering, experience shipping LLM-powered products, evaluation of nondeterministic systems, and ability to diagnose model and system failures.
NVIDIANASDAQ: NVDA: Designs GPU-accelerated computing and artificial intelligence hardware.
12+ YOEMS or PhD in CS/EE/CE, with 12+ years of relevant work; strong impact in large-scale AI training, GPU performance, distributed systems, or similar.
NewsBreak: Local news aggregation platform powered by artificial intelligence.
Hands-on LLM post-training (CPT, SFT, RL) with demonstrated RL experience; strong ML data engineering; experience training LLMs on mid-to-large GPU clusters; PyTorch and related frameworks familiarity; strong communication.
PyTorch, Hugging Face TRL, Hugging Face Accelerate, DeepSpeed, FSDP, vLLM
NebiusNasdaq: NBIS: Builds cloud infrastructure and software for artificial intelligence development.
Expert Python and PyTorch skills, hands-on LLM/VLM inference deployment and optimization, knowledge of modern inference stacks, quantitative reasoning about latency/throughput/cost, and strong communication.
Python, PyTorch, vLLM, SGLang, TensorRT-LLM, Triton Inference Server, NVIDIA Dynamo, Ray Serve, KServe, CUDA, FlashInfer, LMCache, Ray
LLM AIOps Development Engineer - Data Center Networking
San Jose, California, United States
OnsiteFull Time
ByteDance: Developing AI-driven content platforms and mobile applications.
Deep data-center networking and Linux networking knowledge; proficiency in Golang or Python; experience with telemetry, observability, big-data pipelines, and LLM/AIOps approaches; familiarity with protocols like EVPN/VXLAN and BGP/OSPF.
Eightfold: Global AI-native talent intelligence platform provider.
6+ YOELead AI/ML engineer with 6+ years in ML, Gen AI, LLMs; strong Python, TensorFlow/PyTorch; AWS, Docker, Kubernetes; expert in agentic AI and distributed systems.
Lead Machine Learning Engineer - Agentic Models, LLM, RAG, GenAI
Santa Clara, California, United States
$193k-$258k/yrHybridFull Time
Eightfold.ai: AI-native platform for talent management and workforce optimization.
5+ YOESenior ML engineer with expertise in AI agents, LLMs, distributed systems; 5-7+ years of experience; strong Python and ML frameworks; AWS; Docker/Kubernetes; RAG/GenAI experience.
5+ YOEBachelor's in CS/EE/CE or equivalent,5+ years in performance modeling/engineering or architecture,proficiency with C++ or Python,experience with ML serving and hardware/software co-design preferred.
Machine Learning Engineer, Model Evaluations (Speech LLM) - San Francisco
San Francisco, California, United States
$180k-$270k/yrHybridFull Time
Plaud: Develops AI-powered voice recorders and automated transcription software.
Python software engineering; building distributed systems, data pipelines, and evaluation harnesses at scale; partner with ML researchers to define benchmarks; build dashboards and monitor model health; debug mid-training anomalies; communicate results clearly.
Mira Mace: AI-powered healthcare advocacy and navigation for Medicare beneficiaries.
Staff-level backend/full-stack or ML engineering experience, production LLM/ML systems experience, architecture and build-vs-buy judgment, hands-on coding and code review, mentoring and technical leadership.
Baseten: Scalable infrastructure platform for deploying and serving AI models.
4+ YOE1+ MgmtLead a team of Forward Deployed Engineers; strong Python, ML inference, LLM experience; 4+ years software engineering; leadership experience; excellent communication.
Python, vLLM, TensorRT, Triton, Hugging Face, Ray Serve
Shakudo: Develops an operating system for enterprise AI applications.
8+ YOE8+ years engineering experience, 5+ years Kubernetes operation, proficiency in Rust, experience with production infrastructure (physical servers, GPU/DGX clusters), CI/CD, security hardening, observability, and LLM/AI infrastructure.
Beacon AI: Developing an AI-powered-pilot for safer flight operations.
Experience shipping LLM-powered features, building RAG and tool-calling flows, production Python/TypeScript services, and working with embeddings, vector backends, and model providers.