808 llm engineer jobs at 244 companies in Soquel, CA

1mo
Save
Mark Applied
Hide
LLM Inference Engineer
Los Altos, California, United States
OnsiteFull Time
Majestic Labs
Majestic Labs: Developing memory-first AI server platforms for data centers.
3+ YOE3+ years building or operating production LLM inference systems; strong Python and C++; experience with vLLM/SGLang/TensorRT-LLM/Fireworks; distributed inference and performance profiling skills.
vLLM, SGLang, TensorRT-LLM, Fireworks, Python, C++, collective communication library (CCL)
1mo
Save
Mark Applied
Hide
Principal LLM Inference Engineer
Santa Clara, California, United States
$195k-$285k/yr HybridFull Time
d-Matrix: Develops high-performance semiconductor chips for generative AI inference.
10+ YOEBachelor's in CS/EE (or equivalent) with 10+ years experience (Master/PhD with 6+ years preferred); strong Python and C/C++; experience optimizing LLM inference, quantization, batching, GPU kernel programming and contributor-level work on inference frameworks.
Python, C, C++, vLLM, SGLang, TensorRT-LLM, ONNX Runtime, CUDA, Triton, JAX
3mo
Save
Mark Applied
Hide
Distributed LLM Inference Engineer
San Francisco or Palo Alto
$170k-$247k/yr HybridFull Time
Anyscale
Anyscale: Cloud platform for scaling distributed machine learning applications.
Familiarity with running ML inference at large scale with high throughput and low latency; experience with PyTorch; solid understanding of distributed systems.
PyTorch, Ray, vLLM, TensorRT-LLM
3w
Save
Mark Applied
Hide
Senior/Staff LLM Application Engineer - Data Application
San Jose, California, United States
$213k-$450k/yr OnsiteFull Time
TikTok
TikTok: Global short-form video hosting and social media platform.
Experience with data products and LLM application development, strong coding skills in Python, knowledge of prompt engineering, retrieval and benchmarking, and ability to analyze user feedback.
Python
1w
Save
Mark Applied
Hide
Software Engineer, AI Agent & LLM
Mountain View, California, United States
$155k-$185k/yr HybridFull Time
Otter.ai
Otter.ai: AI-powered meeting transcription and automated note-taking platform.
3+ YOE3+ years building AI-agent or ML systems, strong backend/distributed-systems engineering, experience shipping LLM-powered products, evaluation of nondeterministic systems, and ability to diagnose model and system failures.
1mo
Save
Mark Applied
Hide
Machine Learning Engineer, LLM Post-Training
Mountain View, California, United States
$150k-$230k/yr OnsiteFull Time
NewsBreak
NewsBreak: Local news aggregation platform powered by artificial intelligence.
Hands-on LLM post-training (CPT, SFT, RL) with demonstrated RL experience; strong ML data engineering; experience training LLMs on mid-to-large GPU clusters; PyTorch and related frameworks familiarity; strong communication.
PyTorch, Hugging Face TRL, Hugging Face Accelerate, DeepSpeed, FSDP, vLLM
1w
Save
Mark Applied
Hide
Senior Machine Learning Engineer, LLM Inference Optimization
Palo Alto or California
$195k-$262k/yr OnsiteFull Time
Nebius
NebiusNasdaq: NBIS: Builds cloud infrastructure and software for artificial intelligence development.
Expert Python and PyTorch skills, hands-on LLM/VLM inference deployment and optimization, knowledge of modern inference stacks, quantitative reasoning about latency/throughput/cost, and strong communication.
Python, PyTorch, vLLM, SGLang, TensorRT-LLM, Triton Inference Server, NVIDIA Dynamo, Ray Serve, KServe, CUDA, FlashInfer, LMCache, Ray
3w
Save
Mark Applied
Hide
LLM AIOps Development Engineer - Data Center Networking
San Jose, California, United States
OnsiteFull Time
ByteDance
ByteDance: Developing AI-driven content platforms and mobile applications.
Deep data-center networking and Linux networking knowledge; proficiency in Golang or Python; experience with telemetry, observability, big-data pipelines, and LLM/AIOps approaches; familiarity with protocols like EVPN/VXLAN and BGP/OSPF.
gNMI, Netconf, IPFIX, NetFlow, SNMP, Golang, Python, Docker, Kubernetes, CI/CD, Kafka, Flink, ClickHouse, TSDB, Prometheus, OpenTelemetry, Neo4j, SONiC, P4, eBPF, DPDK, RDMA, RoCE
3mo
Save
Mark Applied
Hide
Staff Machine Learning Engineer - Agentic Models, LLM, RAG, GenAI
Santa Clara, California, United States
$232k-$310k/yr HybridFull Time
Eightfold
Eightfold: Global AI-native talent intelligence platform provider.
6+ YOELead AI/ML engineer with 6+ years in ML, Gen AI, LLMs; strong Python, TensorFlow/PyTorch; AWS, Docker, Kubernetes; expert in agentic AI and distributed systems.
Python, TensorFlow, PyTorch, AWS, Docker, Kubernetes
3mo
Save
Mark Applied
Hide
Lead Machine Learning Engineer - Agentic Models, LLM, RAG, GenAI
Santa Clara, California, United States
$193k-$258k/yr HybridFull Time
Eightfold.ai
Eightfold.ai: AI-native platform for talent management and workforce optimization.
5+ YOESenior ML engineer with expertise in AI agents, LLMs, distributed systems; 5-7+ years of experience; strong Python and ML frameworks; AWS; Docker/Kubernetes; RAG/GenAI experience.
Python, TensorFlow, PyTorch, AWS, Docker, Kubernetes, Kafka, AWS SQS, LangGraph, CrewAI, AutoGen, Pinecone, pgvector, vLLM, TensorRT-LLM
4d
Save
Mark Applied
Hide
Senior Performance Co-Design Engineer, LLM Serving
Sunnyvale, California, United States
$174k-$252k/yr OnsiteFull Time
Google
GoogleNASDAQ: GOOGL: Provides online search, advertising, cloud computing, and consumer electronics.
5+ YOEBachelor's in CS/EE/CE or equivalent,5+ years in performance modeling/engineering or architecture,proficiency with C++ or Python,experience with ML serving and hardware/software co-design preferred.
C++, Python, TPU, Vertex AI
2mo
Save
Mark Applied
Hide
Software Engineering Manager, LLM Training
Mountain View, California, United States
$170k-$277k/yr HybridFull Time
LinkedInNASDAQ: MSFT: Professional social network for career development and job recruitment.
1+ Mgmt5+ years software engineering; 1+ year management; experience with LLMs and distributed systems; strong leadership and strategic planning
PyTorch, CUDA, Megatron, Hugging Face, vLLM, Liger, Ray, SGLang, VERL, NCCL, TorchDistributed
1mo
Save
Mark Applied
Hide
Engineering Manager, LLM Performance
Santa Clara, California, United States
$224k-$431k/yr HybridFull Time
NVIDIA
NVIDIANASDAQ: NVDA: Designs GPU-accelerated computing and artificial intelligence hardware.
7+ YOE3+ MgmtMS/PhD or equivalent experience in CS/CE/AI, 7+ years software engineering experience including 3+ years technical leadership; strong C++ or Python; expertise in LLM/VLM/inference and production-quality software.
TensorRT LLM, vLLM, SGLang, Dynamo, C++, Python, CUDA
3d
Save
Mark Applied
Hide
Software Engineer, Artificial Intelligence/LLM (Multiple Seniority Levels)
San Carlos, California, United States
$135k-$260k/yr HybridFull Time
Beacon AI
Beacon AI: Developing an AI-powered-pilot for safer flight operations.
Experience building LLM-powered features, RAG and tool-calling, production services in Python or TypeScript, vector search and embeddings, evals/metrics, and safety/compliance for a regulated domain.
LangChain, Python, TypeScript, AWS Bedrock, OpenAI, Anthropic, OpenSearch, pgvector, Pinecone, Weaviate, S3, Aurora, DynamoDB, Triton, TensorRT-LLM
1mo
Save
Mark Applied
Hide
Infrastructure Engineer
Menlo Park, California, United States
OnsiteFull Time
Shakudo
Shakudo: Develops an operating system for enterprise AI applications.
8+ YOE8+ years engineering experience, 5+ years Kubernetes operation, proficiency in Rust, experience with production infrastructure (physical servers, GPU/DGX clusters), CI/CD, security hardening, observability, and LLM/AI infrastructure.
Kubernetes, Rust, CI/CD, DGX, GPU, LLM, ETL
3w
Save
Mark Applied
Hide
Software Engineer 4 - LLM Inference
San Jose or Durham or Mexico or Canada or India or Netherlands or Serbia or Spain or Singapore or Australia or Japan
$148k-$222k/yr HybridFull Time
Nutanix
NutanixNASDAQ: NTNX: Sells cloud software and hyperconverged infrastructure for enterprises.
6+ YOE6+ years building distributed, high-performance cloud-native systems; strong Go/Python, Docker, Kubernetes, CI/CD, systems and networking knowledge; bachelor’s or master’s in CS or equivalent.
Docker, Kubernetes, Go, Python, CI/CD, TensorFlow, PyTorch
2mo
Save
Mark Applied
Hide
Lead AI Engineer (AI Foundations, LLM Customization and Finetuning)
Cambridge or McLean or San Jose or New York City
$197k-$225k/yr OnsiteFull Time
Capital One
Capital OneNYSE: COF: Financial services offering credit cards, banking, and loans.
4+ YOEBachelor's in CS/AI/EE/CE with 4+ years or Master's with 2+ years in AI/ML; 4+ years Python/Go/Scala/Java
AWS, Hugging Face, VectorDBs, Nemo Guardrails, PyTorch, Python, Go, Scala, Java, C++
1mo
Save
Mark Applied
Hide
Staff Design Engineer, Insights
Minnesota or San Francisco or San Jose
$159k-$302k/yr RemoteFull Time
Adobe
AdobeNASDAQ: ADBE: Provides software for digital media creation and marketing analytics
10+ YOE10+ years at the intersection of design and engineering; hands-on experience building and shipping LLM-powered apps, internal tools, and high-fidelity prototypes; strong craft, collaboration, and product judgment.
LLM, Vega, D3, Adobe Acrobat Studio, Adobe Express, Adobe Firefly, Creative Cloud, Adobe Experience Platform, Adobe Experience Manager, GenStudio
2w
Save
Mark Applied
Hide
Staff Software Engineer, Agentic Applications
Mountain View, California, United States
$198k-$273k/yr OnsiteFull Time
Databricks
Databricks: A unified platform for data analytics and artificial intelligence.
12+ YOE12+ years software engineering experience, production LLM systems experience, human-in-the-loop design, CMS and web publishing pipeline knowledge, stakeholder collaboration, and mentoring experience.
Apache Spark, Delta Lake, MLflow, CMS, LLM
1mo
Save
Mark Applied
Hide
Staff Engineer, Agents
Santa Clara, California, United States
$160k-$200k/yr HybridFull Time
LeanData
LeanData: Automates lead-to-account matching and routing for B2B revenue teams.
4+ YOE4+ years building production systems (2+ years shipping LLM/agent systems), strong Python (TypeScript/Go a plus), experience with agent frameworks, RAG/memory, distributed-systems fundamentals, and shipping customer-facing LLM products.
Python, TypeScript, Go, LangGraph, agent SDKs, LLM, RAG, Salesforce APIs, Bulk 2.0, Composite, Pub/Sub, MCP (Model Context Protocol), Promptfoo, Braintrust, Langfuse, Inngest, Temporal, Postgres, RLS