35 inference infrastructure software engineer jobs at 24 companies in San Francisco, CA
3mo
Save
Mark Applied
Hide
3mo
INFERENCE ENGINEER
San Francisco, California, United States
OnsiteFull Time
MakerMaker: Small San Francis-based AI startup focused on autonomous agents and production ML systems.
3+ YOESenior ML systems engineer with 3+ years building production-grade, large-scale serving infrastructure; strong distributed systems; GPU-accelerated inference; fluent Python and systems languages (C++, CUDA, ROCm or Triton).
Senior Software Engineer, AI Infrastructure - LVM Inference & Evaluation
Redwood City, California, United States
$168k-$205k/yrHybridFull Time
Ambient.ai: AI-powered physical security platform for proactive threat detection.
4+ YOE4+ years building infrastructure or production AI systems; strong Python; experience with ML infrastructure, LLM/LVM inference, inference optimization, evaluation frameworks, cloud and GPU workloads; BS/MS or equivalent.
Nuro: Builds autonomous driving software and electric delivery robots.
3+ YOE3+ years of software engineering experience, or 2+ with a Master's; strong Python, LLM research, agent systems, backend and distributed systems, ML infrastructure, evaluation, and inference experience.
Cerebras SystemsNasdaq: CBRS: Manufactures specialized computer chips designed for AI.
8+ YOE8+ years software engineering experience building and operating large-scale distributed systems or cloud infrastructure; deep distributed systems and Kubernetes expertise; proficiency in Go or C++; security and observability experience.
Anthropic: Developing safe and reliable artificial intelligence systems.
Senior IC with deep systems or ML infrastructure experience, hands-on performance profiling and optimization, accelerator ecosystem expertise (CUDA/TPU/Trainium), strong software engineering and cross-org alignment skills, and a relevant bachelor’s degree or equivalent.
ByteDance: Developing AI-driven content platforms and mobile applications.
2+ YOEB.S./M.S. in CS or CE with 2+ years experience; strong knowledge of large-model inference, distributed systems, container orchestration; proficiency in Go, Rust, Python, or C++; experience with Kubernetes and cloud/ML infrastructure.
Baseten: Scalable infrastructure platform for deploying and serving AI models.
Bachelor’s, Master’s, or Ph.D. in Computer Science, Engineering, or a related field; strong background in distributed systems and backend infrastructure; production experience; excellent communication and collaboration skills.
Senior Software Engineer, Machine Learning Infrastructure - Generative AI
San Francisco or Sunnyvale or Seattle
$137k-$299k/yrOnsiteFull Time
DoorDashNYSE: DASH: Local food delivery and on-demand logistics platform.
6+ YOEB.S., M.S., or PhD in Computer Science or equivalent; 6+ years software engineering experience; Python, distributed systems, production ML infrastructure, LLM inference or fine-tuning, technical leadership, and AI coding tools.
Principal Software Engineer, Machine Learning Infrastructure
Palo Alto or Seattle or Los Angeles or New York or Bellevue
$235k-$414k/yrOnsiteFull Time
SnapNYSE: SNAP: Develops social media applications and augmented reality technology.
10+ YOE10+ years software development experience, technical leadership, distributed systems and ML inference platform expertise, strong software design and debugging skills, ability to operate highly-available systems at scale.
8+ YOE3+ MgmtBachelor's degree or equivalent,8+ years software development,5+ years testing/launching and ML/ML infrastructure experience,3+ years software design/architecture and technical leadership; experience with TPUs/GPUs and ML runtimes.
Senior Software Engineer, Machine Learning Infrastructure - Generative AI
San Francisco or Sunnyvale or Seattle
$137k-$202k/yrOnsiteFull Time
DoorDashNASDAQ: DASH: On-demand delivery platform connecting consumers with local merchants.
6+ YOE6+ years software engineering experience; BS/MS/PhD in CS or equivalent; deep backend fundamentals in Python and distributed systems; experience with LLM inference/fine-tuning, production reliability, observability, and technical leadership.
NIONYSE: NIO: Designs and manufactures premium smart electric vehicles and technology
5+ YOE5+ years building and optimizing large-scale LLM/VLM inference systems; strong C/C++ and performance engineering skills; GPU/NPU programming (CUDA), PyTorch/TensorFlow, and BS/MS in CS/CE or related field required.
Member of Technical Staff - Machine Learning Infrastructure Engineer
San Francisco or Toronto or Seattle
$180k-$300k/yrOnsiteFull Time
Preference Model: Building reinforcement learning environments to train frontier AI models.
Experienced software engineer with production ML/data infrastructure skills, proficiency with PyTorch or JAX, distributed systems, AWS/GCP, Kubernetes, data pipelines, and familiarity with transformers and inference libraries like vLLM.
NVIDIANASDAQ: NVDA: Designs GPU-accelerated computing and artificial intelligence hardware.
12+ YOE12+ years software engineering experience in systems software, AI/ML infrastructure, deep learning inference, or compiler/runtime; strong C/C++ and Python; familiarity with TensorRT, ONNX, PyTorch, CUDA, and model optimization.
Sr. Machine Learning Engineer, Foundation Models Inference - Cloud OS & Inference
Santa Clara, California, United States
OnsiteFull Time
AppleNASDAQ: AAPL: Designs and sells consumer electronics, software, and online services.
Build and optimize inference frameworks, services, and tools for large-scale foundation models across cloud infrastructure, including language, vision, and speech models.
Rippling: Unified platform managing workforce HR, IT, and finance operations
8+ YOE8+ years software engineering experience, distributed systems ownership, experience training/deploying LLMs, model inference optimization, backend skills in Python/Go/Java, and cloud-native infrastructure (Kubernetes).
TikTok: Global short-form video hosting and social media platform.
3+ YOE3+ years building scalable ML systems, strong CS fundamentals, coding skills, experience with causal inference/uplift/deep learning, project management and communication skills.
Patreon: Membership platform for creators to monetize work and engage fans.
Deep experience building production ML infrastructure, low-latency inference pipelines, feature stores, distributed systems, backend engineering, Python, debugging, documentation, and code reviews.
Engineering Manager, Inference Benchmarking — AI Perf
Santa Clara or Alabama or Austin or United States or Washington or California
$224k-$357k/yrHybridFull Time
NVIDIANASDAQ: NVDA: Designs graphics processing units and artificial intelligence hardware.
8+ YOE3+ Mgmt8+ years software engineering experience, 3+ years engineering leadership, strong systems and inference infrastructure expertise, deep understanding of LLM inference mechanics, and ability to deliver production-quality output in high-visibility environments.