35 inference infrastructure software engineer jobs at 24 companies in San Francisco, CA

3mo
Save
Mark Applied
Hide
INFERENCE ENGINEER
San Francisco, California, United States
OnsiteFull Time
MakerMaker
MakerMaker: Small San Francis-based AI startup focused on autonomous agents and production ML systems.
3+ YOESenior ML systems engineer with 3+ years building production-grade, large-scale serving infrastructure; strong distributed systems; GPU-accelerated inference; fluent Python and systems languages (C++, CUDA, ROCm or Triton).
Python, C++, CUDA, ROCm, Triton
1mo
Save
Mark Applied
Hide
Senior Software Engineer, AI Infrastructure - LVM Inference & Evaluation
Redwood City, California, United States
$168k-$205k/yr HybridFull Time
Ambient.ai
Ambient.ai: AI-powered physical security platform for proactive threat detection.
4+ YOE4+ years building infrastructure or production AI systems; strong Python; experience with ML infrastructure, LLM/LVM inference, inference optimization, evaluation frameworks, cloud and GPU workloads; BS/MS or equivalent.
Python, vLLM, Triton Inference Server, CUDA, NCCL, PyTorch, TensorRT, ONNX
1w
Save
Mark Applied
Hide
Software Engineer, Applied AI Infrastructure
Mountain View, California, United States
$194k-$352k/yr OnsiteFull Time
Nuro
Nuro: Builds autonomous driving software and electric delivery robots.
3+ YOE3+ years of software engineering experience, or 2+ with a Master's; strong Python, LLM research, agent systems, backend and distributed systems, ML infrastructure, evaluation, and inference experience.
Python, Go, C++, Rust, SFT, RL
2mo
Save
Mark Applied
Hide
Staff Software Engineer, Inference Platform
Sunnyvale or Toronto
OnsiteFull Time
Cerebras Systems
Cerebras SystemsNasdaq: CBRS: Manufactures specialized computer chips designed for AI.
8+ YOE8+ years software engineering experience building and operating large-scale distributed systems or cloud infrastructure; deep distributed systems and Kubernetes expertise; proficiency in Go or C++; security and observability experience.
Kubernetes, Go, C++, TLS, mTLS
2mo
Save
Mark Applied
Hide
Staff+ Software Engineer, Inference Runtime
San Francisco or Seattle or New York City
$405k-$485k/yr HybridFull Time
Anthropic
Anthropic: Developing safe and reliable artificial intelligence systems.
Senior IC with deep systems or ML infrastructure experience, hands-on performance profiling and optimization, accelerator ecosystem expertise (CUDA/TPU/Trainium), strong software engineering and cross-org alignment skills, and a relevant bachelor’s degree or equivalent.
Rust, Python, CUDA, XLA, Triton, NeuronX, AWS Neuron, Kubernetes, CI/CD
3w
Save
Mark Applied
Hide
Staff Software Engineer- Foundation Model Inference
San Francisco or Mountain View
$190k-$265k/yr OnsiteFull Time
Databricks
Databricks: A unified platform for data analytics and artificial intelligence.
8+ YOE8+ years backend or infrastructure engineering experience; distributed systems, scalable APIs, real-time serving or ML/GPU orchestration experience; familiarity with service-oriented architecture, deployment pipelines, and observability.
OpenAI, Anthropic, Gemini, Qwen, GPT-OSS, Llama, SageMaker, Vertex AI, Azure ML, MLflow, PyTorch, Ray, vLLM, SGLang, Apache Spark, Delta Lake
1mo
Save
Mark Applied
Hide
Software Engineer - AI Compute Infrastructure
San Jose, California, United States
OnsiteFull Time
ByteDance
ByteDance: Developing AI-driven content platforms and mobile applications.
2+ YOEB.S./M.S. in CS or CE with 2+ years experience; strong knowledge of large-model inference, distributed systems, container orchestration; proficiency in Go, Rust, Python, or C++; experience with Kubernetes and cloud/ML infrastructure.
AIBrix, vLLM, SGLang, TensorRT-LLM, Kubernetes, Docker, Ray, CUDA, AWS, Azure, GCP, SageMaker, Azure ML, Vertex AI, DeepSpeed, PyTorch, Go, Rust, Python, C++
2mo
Save
Mark Applied
Hide
Software Engineer- BIS (Baseten Inference Stack)
San Francisco, California, United States
$180k-$360k/yr HybridFull Time
Baseten
Baseten: Scalable infrastructure platform for deploying and serving AI models.
Bachelor’s, Master’s, or Ph.D. in Computer Science, Engineering, or a related field; strong background in distributed systems and backend infrastructure; production experience; excellent communication and collaboration skills.
2w
Save
Mark Applied
Hide
Senior Software Engineer, Machine Learning Infrastructure - Generative AI
San Francisco or Sunnyvale or Seattle
$137k-$299k/yr OnsiteFull Time
DoorDash
DoorDashNYSE: DASH: Local food delivery and on-demand logistics platform.
6+ YOEB.S., M.S., or PhD in Computer Science or equivalent; 6+ years software engineering experience; Python, distributed systems, production ML infrastructure, LLM inference or fine-tuning, technical leadership, and AI coding tools.
Python, LLM, VLM, GLM, Qwen, Kimi, DeepSeek, SFT, DPO, LoRA, RLHF, RLVR, Claude Code, Codex, Cursor, vLLM, SGLang, TensorRT-LLM, Kubernetes, AWS, GCP, Modal, MCP, RAG, FP8, INT8, AWQ, GPTQ, Gem, Covey
1mo
Save
Mark Applied
Hide
Principal Software Engineer, Machine Learning Infrastructure
Palo Alto or Seattle or Los Angeles or New York or Bellevue
$235k-$414k/yr OnsiteFull Time
Snap
SnapNYSE: SNAP: Develops social media applications and augmented reality technology.
10+ YOE10+ years software development experience, technical leadership, distributed systems and ML inference platform expertise, strong software design and debugging skills, ability to operate highly-available systems at scale.
Tensorflow, PyTorch, Kubernetes, GPU, LLM inference, RPC
2w
Save
Mark Applied
Hide
Staff Software Engineer, AI Engines 3P TPU Inference
Mountain View, California, United States
$207k-$300k/yr OnsiteFull Time
Google
GoogleNASDAQ: GOOGL: Provides online search, advertising, cloud computing, and consumer electronics.
8+ YOE3+ MgmtBachelor's degree or equivalent,8+ years software development,5+ years testing/launching and ML/ML infrastructure experience,3+ years software design/architecture and technical leadership; experience with TPUs/GPUs and ML runtimes.
TPUs, GPUs, TensorFlow, JAX, PyTorch, TF Executor, TFRT, PJRT
1mo
Save
Mark Applied
Hide
Senior Software Engineer, Machine Learning Infrastructure - Generative AI
San Francisco or Sunnyvale or Seattle
$137k-$202k/yr OnsiteFull Time
DoorDash
DoorDashNASDAQ: DASH: On-demand delivery platform connecting consumers with local merchants.
6+ YOE6+ years software engineering experience; BS/MS/PhD in CS or equivalent; deep backend fundamentals in Python and distributed systems; experience with LLM inference/fine-tuning, production reliability, observability, and technical leadership.
Python, Claude Code, Codex, Cursor, vLLM, SGLang, TensorRT-LLM, Kubernetes, AWS, GCP, Modal
2mo
Save
Mark Applied
Hide
AI Infrastructure Engineer
San Jose, California, United States
$192k-$250k/yr OnsiteFull Time
NIO
NIONYSE: NIO: Designs and manufactures premium smart electric vehicles and technology
5+ YOE5+ years building and optimizing large-scale LLM/VLM inference systems; strong C/C++ and performance engineering skills; GPU/NPU programming (CUDA), PyTorch/TensorFlow, and BS/MS in CS/CE or related field required.
CUDA, PyTorch, TensorFlow, C/C++, AIOS
1mo
Save
Mark Applied
Hide
Member of Technical Staff - Machine Learning Infrastructure Engineer
San Francisco or Toronto or Seattle
$180k-$300k/yr OnsiteFull Time
Preference Model
Preference Model: Building reinforcement learning environments to train frontier AI models.
Experienced software engineer with production ML/data infrastructure skills, proficiency with PyTorch or JAX, distributed systems, AWS/GCP, Kubernetes, data pipelines, and familiarity with transformers and inference libraries like vLLM.
PyTorch, JAX, AWS, GCP, Kubernetes, transformers, vLLM, SGLang
1mo
Save
Mark Applied
Hide
Senior Software Engineer - Autonomous Driving
Santa Clara, California, United States
$224k-$357k/yr OnsiteFull Time
NVIDIA
NVIDIANASDAQ: NVDA: Designs GPU-accelerated computing and artificial intelligence hardware.
12+ YOE12+ years software engineering experience in systems software, AI/ML infrastructure, deep learning inference, or compiler/runtime; strong C/C++ and Python; familiarity with TensorRT, ONNX, PyTorch, CUDA, and model optimization.
C/C++, Python, TensorRT, TensorRT-LLM, ONNX, PyTorch, CUDA, Triton, DriveOS, QNX, Safe RTOS, Linux, hypervisors
5d
Save
Mark Applied
Hide
Sr. Machine Learning Engineer, Foundation Models Inference - Cloud OS & Inference
Santa Clara, California, United States
OnsiteFull Time
Apple
AppleNASDAQ: AAPL: Designs and sells consumer electronics, software, and online services.
Build and optimize inference frameworks, services, and tools for large-scale foundation models across cloud infrastructure, including language, vision, and speech models.
2w
Save
Mark Applied
Hide
Staff Software Engineer - Data Cloud Applied ML
San Francisco or Seattle or New York City
$189k-$315k/yr OnsiteFull Time
Rippling
Rippling: Unified platform managing workforce HR, IT, and finance operations
8+ YOE8+ years software engineering experience, distributed systems ownership, experience training/deploying LLMs, model inference optimization, backend skills in Python/Go/Java, and cloud-native infrastructure (Kubernetes).
Python, Go, Java, Kubernetes
1mo
Save
Mark Applied
Hide
Software Engineer, Ads ML Infrastructure
San Jose, California, United States
$156k-$317k/yr OnsiteFull Time
TikTok
TikTok: Global short-form video hosting and social media platform.
3+ YOE3+ years building scalable ML systems, strong CS fundamentals, coding skills, experience with causal inference/uplift/deep learning, project management and communication skills.
1d
Save
Mark Applied
Hide
Senior Machine Learning Engineer, Infrastructure
New York City or San Francisco
$212k-$318k/yr HybridFull Time
Patreon
Patreon: Membership platform for creators to monetize work and engage fans.
Deep experience building production ML infrastructure, low-latency inference pipelines, feature stores, distributed systems, backend engineering, Python, debugging, documentation, and code reviews.
Python
2mo
Save
Mark Applied
Hide
Engineering Manager, Inference Benchmarking — AI Perf
Santa Clara or Alabama or Austin or United States or Washington or California
$224k-$357k/yr HybridFull Time
NVIDIA
NVIDIANASDAQ: NVDA: Designs graphics processing units and artificial intelligence hardware.
8+ YOE3+ Mgmt8+ years software engineering experience, 3+ years engineering leadership, strong systems and inference infrastructure expertise, deep understanding of LLM inference mechanics, and ability to deliver production-quality output in high-visibility environments.
AIPerf, vLLM, TRT-LLM, SGLang, Kubernetes, Helm, DCGM, dcgm-exporter, PyNVML, Prometheus, ZMQ, MLPerf