218 ai ml infrastructure engineer jobs at 88 companies in Berkeley, CA

PromotedHiringCafe
ML Engineer - Inference & Model Deployment
Cupertino, CA, US
$250k-$310k/yr On-SiteFull Time
HiringCafe
HiringCafe: Building a 100× better job search engine to take on Indeed and LinkedIn.
Turn powerful AI and ML models into fast, reliable production systems. Own inference latency, throughput, model-serving architecture, multi-GPU systems, and production deployment for millions of users.
Python, PyTorch, vLLM, SGLang, TensorRT, LLMs
PromotedHiringCafe
Founding Machine Learning / AI Search Engineer
Cupertino, CA, US
$160k-$310k/yr On-SiteFull Time
HiringCafe
HiringCafe: Building a 100× better job search engine to take on Indeed and LinkedIn.
Build the ML and AI search behind HiringCafe — ranking, recommenders, retrieval, and LLM agents that surface jobs people would never find on their own.
Python, PyTorch, Elasticsearch, LLMs
PromotedHiringCafe
Founding Backend / Infra Engineer
Cupertino, CA, US
$160k-$300k/yr On-SiteFull Time
HiringCafe
HiringCafe: Building a 100× better job search engine to take on Indeed and LinkedIn.
Own the crawlers, pipelines, and infrastructure powering a real-time job search engine. Strong Node.js and Python fundamentals; bonus points for security and reverse-engineering chops.
Node.js, Python, Elasticsearch, Redis
1mo
Save
Mark Applied
Hide
AI Infrastructure Engineer
San Jose, California, United States
$179k-$306k/yr HybridFull Time
AMD
AMDNASDAQ: AMD: Designs and manufactures computer processors and graphics technology.
5+ YOE5+ years in DevOps/platform/infrastructure engineering; deep Kubernetes experience; experience building developer-facing platforms, Helm and GitOps workflows, storage/networking for GPU workloads; Terraform, monitoring, and ML framework exposure preferred.
Kubernetes, Helm, ArgoCD, Flux, GitOps, Terraform, CSI drivers, CNI, Prometheus, Grafana, Loki, Slurm, PyTorch, vLLM, SGLang
1mo
Save
Mark Applied
Hide
Senior AI ML Engineer - Remote
San Francisco or Minneapolis or Washington, D.C. or United States
$120k-$215k/yr RemoteFull Time
UnitedHealth Group
UnitedHealth GroupNYSE: UNH: Provides health insurance and technology-enabled health care services.
4+ YOE2+ MgmtBachelor's degree or 4+ years equivalent, 4+ years Python, 4+ years cloud infrastructure (AWS/Azure/GCP), 4+ years AI/ML infrastructure experience, 2+ years team lead, 1+ year LLM experience.
Python, AWS, Azure, GCP, Large Language Models (LLMs), GitHub, GitHub Actions, Docker, Terraform, CI/CD
3mo
Save
Mark Applied
Hide
Senior ML Infrastructure Engineer, Proactive
Cupertino or Seattle
OnsiteFull Time
Apple
AppleNASDAQ: AAPL: Designs and sells consumer electronics, software, and online services.
Experience designing large-scale ML/AI platforms; proficiency in ML systems, LLMs, and distributed data; strong multi-threaded programming skills.
Python, Distributed systems, MacOS/iOS development, LLMs, Machine learning platforms
1mo
Save
Mark Applied
Hide
Senior AI/ML Platform Engineer
San Mateo, California, United States
$148k-$247k/yr HybridFull Time
Guidewire
GuidewireNYSE: GWRE: Provides a software platform for property and casualty insurers.
10+ YOE10+ years software engineering; 5+ years ML platforms/infrastructure; distributed systems; Python/Go/Java; Docker/Kubernetes; MLOps tools; cloud experience; knowledge of ML models.
Python, Go, Java, Docker, Kubernetes, MLflow, Kubeflow, SageMaker, Vertex AI, Databricks, AWS, GCP, Azure
1mo
Save
Mark Applied
Hide
Senior AI / ML Engineer
Melbourne or San Francisco or Sydney
OnsiteFull Time
Artificial Analysis
Artificial Analysis: Independent AI benchmarking and performance analysis platform.
3+ YOE3+ years software engineering experience; strong Python, pandas and OpenAI API proficiency; experience with data-intensive backend systems, data visualization, cloud infrastructure, and leading projects end-to-end.
Python, pandas, OpenAI, PyTorch
1mo
Save
Mark Applied
Hide
ML Engineer
Palo Alto or Seattle or Paris
$139k-$226k/yr RemoteFull Time
Docker
Docker: Provides a platform for building, sharing, and running containerized applications.
5+ YOE5+ years applied ML/AI experience, 4+ years software engineering, experience with LLM-based systems, model lifecycle and ML infrastructure, bachelor's in CS/Engineering or equivalent, strong communication and mentoring skills.
Docker Desktop, Docker Hub, Docker Scout, LLM, MCP, Agentic Platform
1mo
Save
Mark Applied
Hide
Senior Applied AI and AI Infrastructure Engineer - Chip Design and DFX
Santa Clara, California, United States
$200k-$380k/yr OnsiteFull Time
NVIDIA
NVIDIANASDAQ: NVDA: Designs graphics processing units and artificial intelligence hardware.
6+ YOEAdvanced degree or equivalent experience in engineering with 6+–12+ yrs experience in AI infrastructure, applied ML/GenAI; experience with agents, SQL/ETL/data modeling, cloud (AWS/Azure/GCP), Python and C++; strong communication.
SQL, ETL, AWS, Azure, GCP, Python, C++
1mo
Save
Mark Applied
Hide
Staff AI/ML Software Engineer, YouTube Ads Creative Foundational Infrastructure
Mountain View, California, United States
$207k-$301k/yr OnsiteFull Time
Google
GoogleNASDAQ: GOOGL: Provides online search, advertising, cloud computing, and consumer electronics.
8+ YOEBachelor's in CS or equivalent, 8+ years software development, 5+ years building large-scale infrastructure, 3+ years software design, experience with ML infrastructure; leadership and deep learning framework experience preferred.
TensorFlow, PyTorch
1mo
Save
Mark Applied
Hide
Senior Applied AI and AI Infrastructure Engineer - Chip Design and DFX
Santa Clara, California, United States
$200k-$380k/yr OnsiteFull Time
NVIDIA
NVIDIANASDAQ: NVDA: Designs GPU-accelerated computing and artificial intelligence hardware.
6+ YOEBSEE/MSEE/PhD with 12+/10+/6+ years experience in AI infrastructure, applied ML and Gen AI. Experience with SQL, ETL, data modeling, cloud (AWS/Azure/GCP), Python, C++, distributed systems, and mentoring.
SQL, ETL, Python, C++, AWS, Azure, GCP
2w
Save
Mark Applied
Hide
Senior AI Infrastructure Engineer - Model Training
Mountain View, California, United States
$190k-$260k/yr OnsiteFull Time
Kodiak Robotics
Kodiak RoboticsNASDAQ: KDK: Develops autonomous driving technology for commercial trucking and defense.
2+ YOEDegree in CS or related field,2+ years ML systems experience,expertise in distributed training,high-performance data pipelines,GPU performance and profiling,Python and PyTorch skills.
PyTorch, PyTorch DDP/FSDP, DeepSpeed, Megatron, NCCL, WebDataset, MosaicML Streaming, MDS, Nsight, PyTorch Profiler, Python, C++, CUDA, Triton, NVLink, InfiniBand
3mo
Save
Mark Applied
Hide
Senior AI/ML Capacity and Performance Engineer
Sunnyvale or Seattle
$145k-$261k/yr HybridFull Time
General Motors
General MotorsNYSE: GM: Manufactures and sells automobiles and automotive parts globally.
5+ YOE5+ years in high-scale infrastructure or ML systems; BS in Computer Science or related field; strong Python and PyTorch; Kubernetes; cloud experience; GPU/ML infra focus.
Python, PyTorch, Kubernetes, Nvidia DCGM, nvidia-smi, Grafana, AWS, GCP, Azure, BigQuery, Hugging Face, Nvidia Nsight, Nsight Compute
2mo
Save
Mark Applied
Hide
Software Engineer, AI/ML (Infrastructure & Platform)
New York City or Tempe or San Francisco
$200k-$235k/yr HybridFull Time
Wealth.com
Wealth.com: Digital estate planning platform for wealth management professionals.
Build AI infrastructure and platform software; strong software engineering, distributed systems, production AI experience.
Python, TypeScript, C#, APIs, Distributed Systems, LLM/AI Systems, Tool/Skill Abstractions, LangChain, Data Pipelines
3w
Save
Mark Applied
Hide
Senior MLOps & AI Infrastructure Engineer
San Jose, California, United States
$149k-$216k/yr OnsiteFull Time
Altera
Altera: Manufacturer of field-programmable gate arrays and programmable logic devices.
10+ YOEBachelor's degree, 10+ years ML engineering/MLOps experience, strong Python, cloud ML platforms (AWS/GCP/Azure), Docker/Kubernetes, CI/CD, ML frameworks (PyTorch/TensorFlow/JAX), experience with MLflow/W&B and HPC schedulers.
PyTorch, TensorFlow, JAX, Hugging Face, scikit-learn, XGBoost, MLflow, Kubeflow, Airflow, Weights & Biases, DVC, Feast, AWS SageMaker, GCP Vertex AI, Azure ML, Terraform, CloudFormation, Docker, Kubernetes, Slurm, LSF, Python, Bash, Go, SQL, Prometheus, Grafana, ELK Stack, Evidently AI, Arize
3mo
Save
Mark Applied
Hide
Principal Engineer, AI Platform & Infrastructure
San Francisco, California, United States
HybridFull Time
SpreeAI
SpreeAI: AI-powered virtual try-on and sizing software for fashion retailers.
10+ YOE10+ years software engineering/infrastructure; 5+ years ML infrastructure, MLOps, or AI platform engineering; strong Python, PyTorch, Kubernetes, Docker; distributed systems expertise; experience with ML workflow orchestration and production inference systems.
Python, PyTorch, Kubernetes, Docker, cloud infrastructure, GPU workloads
2w
Save
Mark Applied
Hide
Sr. Mgr., ML Infrastructure, PV Personalization and Discovery
Sunnyvale or Seattle or New York
$242k-$328k/yr OnsiteFull Time
Amazon
AmazonNASDAQ: AMZN: Global online retail and cloud computing technology provider.
10+ YOE5+ Mgmt10+ years engineering experience, 5+ years managing engineering teams, expertise in ML infrastructure, retrieval/recommendation systems, LLMs/generative AI, partnering with applied scientists, experience delivering large-scale consumer software.
AWS
2mo
Save
Mark Applied
Hide
Infrastructure Engineer
New York or San Francisco or United States
$165k-$200k/yr HybridFull Time
Roboflow
Roboflow: Platform for building and deploying custom computer vision models.
Kubernetes production experience; IaC (Terraform/Helm); cloud (AWS/GCP); Python/Node.js; CI/CD (GitHub Actions/Spacelift); security and ML/AI infrastructure familiarity.
Kubernetes, Terraform, Helm, Python, Node.js, GitHub Actions, Spacelift, AWS, GCP, PyTorch, TensorFlow, Bash
1mo
Save
Mark Applied
Hide
Founding Engineer, AI Infra
San Francisco, California, United States
HybridFull Time
Phenix Space
Phenix Space: Builds robotic systems for on-orbit satellite assembly and upgrades.
5+ YOE5+ years building or operating ML infrastructure; deep GPU and distributed training knowledge; experience with PyTorch/DeepSpeed/Megatron/Ray, inference stacks (vLLM, TGI, Triton), Python and C++/Rust/Go, Kubernetes and IaC, and observability tooling.
FlashAttention, CUDA, Triton, PyTorch, DeepSpeed, Megatron, Ray, vLLM, SGLang, TGI, Python, C++, Rust, Go, Kubernetes, Terraform, Pulumi, Prometheus, Grafana, OpenTelemetry, Llama 3, Qwen, DeepSeek
3mo
Save
Mark Applied
Hide
Hyperbolic Labs - Senior GPU Infrastructure Engineer
San Francisco, California, United States
RemoteFull Time
YieldNest
YieldNest: Liquid restaking protocol for risk-adjusted DeFi yields.
Senior infrastructure/DevOps engineer with expertise in bare-metal provisioning, GPU scheduling, Terraform/Pulumi, CI/CD for infrastructure, storage for AI/ML workloads, and cloud-init provisioning.
Terraform, Pulumi, CI/CD, infrastructure as code, secrets management, configuration management, observability stack, object storage, block storage, distributed file systems, cloud-init, CUDA, GPU topology, GPU orchestration
3mo
Save
Mark Applied
Hide
Forward Deployed AI Engineer
San Francisco, California, United States
HybridFull Time
Latent Labs
Latent Labs: Building generative AI models to make biology programmable.
Strong CS/ML background; experience deploying AI systems; cloud infrastructure; customer-facing; production software; pharma/biotech domain knowledge.
AWS, GCP, Azure, Docker, Kubernetes, CI/CD, Cloud-native, APIs, Python, ML frameworks
4d
Save
Mark Applied
Hide
Founding AI Engineer
San Francisco, California, United States
$150k-$200k/yr OnsiteFull Time
Clera
Clera: AI talent agent matching professionals with high-growth startup roles
0+ YOEEligible to work in the US without sponsorship. 0–4 years building production AI/ML systems; experience with collaborative filtering, graph embeddings, reinforcement learning, multi-channel agent design, and scalable AI infrastructure.
Instagram, iMessage, WhatsApp, TikTok