570 ai infrastructure engineer jobs at 237 companies in Berkeley, CA

PromotedHiringCafe
ML Engineer - Inference & Model Deployment
Cupertino, CA, US
$250k-$310k/yr On-SiteFull Time
HiringCafe
HiringCafe: Building a 100× better job search engine to take on Indeed and LinkedIn.
Turn powerful AI and ML models into fast, reliable production systems. Own inference latency, throughput, model-serving architecture, multi-GPU systems, and production deployment for millions of users.
Python, PyTorch, vLLM, SGLang, TensorRT, LLMs
PromotedHiringCafe
Founding Machine Learning / AI Search Engineer
Cupertino, CA, US
$160k-$310k/yr On-SiteFull Time
HiringCafe
HiringCafe: Building a 100× better job search engine to take on Indeed and LinkedIn.
Build the ML and AI search behind HiringCafe — ranking, recommenders, retrieval, and LLM agents that surface jobs people would never find on their own.
Python, PyTorch, Elasticsearch, LLMs
PromotedHiringCafe
Founding Backend / Infra Engineer
Cupertino, CA, US
$160k-$300k/yr On-SiteFull Time
HiringCafe
HiringCafe: Building a 100× better job search engine to take on Indeed and LinkedIn.
Own the crawlers, pipelines, and infrastructure powering a real-time job search engine. Strong Node.js and Python fundamentals; bonus points for security and reverse-engineering chops.
Node.js, Python, Elasticsearch, Redis
1w
Save
Mark Applied
Hide
AI Infrastructure Engineer
San Jose, California, United States
$45k-$121k/yr OnsiteFull Time
Wipro
WiproNYSE: WIT: Global technology services and consulting for digital transformation.
3+ YOEDesign and build hybrid AI infrastructure integrating compute, GPU nodes, software-defined storage, Kubernetes, IaC, and observability; 3+ years experience in AI infrastructure or related roles.
NC2, Objects (S3-compatible), Kubernetes, Calm, Terraform, Prometheus, Grafana, ELK, OpenTelemetry, AOS, AHV, Cloud Manager (NCM)
2mo
Save
Mark Applied
Hide
AI Infrastructure Engineer
San Francisco, California, United States
$190k-$270k/yr OnsiteFull Time
Together AI
Together AI: Cloud platform for training and deploying artificial intelligence models.
5+ YOE5+ years in AI infrastructure or related roles; BS in CS or equivalent; knowledge of Ansible, Terraform, Kubernetes; programming/scripting; monitoring/observability; cloud services; collaborative work
Ansible, Terraform, Kubernetes
1mo
Save
Mark Applied
Hide
AI Infrastructure Engineer
San Jose, California, United States
$179k-$306k/yr HybridFull Time
AMD
AMDNASDAQ: AMD: Designs and manufactures computer processors and graphics technology.
5+ YOE5+ years in DevOps/platform/infrastructure engineering; deep Kubernetes experience; experience building developer-facing platforms, Helm and GitOps workflows, storage/networking for GPU workloads; Terraform, monitoring, and ML framework exposure preferred.
Kubernetes, Helm, ArgoCD, Flux, GitOps, Terraform, CSI drivers, CNI, Prometheus, Grafana, Loki, Slurm, PyTorch, vLLM, SGLang
6d
Save
Mark Applied
Hide
AI Infrastructure Engineer
Fremont, California, United States
OnsiteFull Time
AMAX
AMAXTaiwan Stock Exchange: 6933: Designs and manufactures GPU-accelerated AI and HPC computing infrastructure.
Experience with on-prem/datacenter operations, Infrastructure-as-Code, networking (VLANs/routing/firewalls), container orchestration, scripting, Git; comfortable with hands-on hardware tasks.
Terraform, Terragrunt, Ansible, Vault, Boundary, Keycloak, Prometheus, Grafana, Alertmanager, Docker, Kubernetes, Git, Jira, Confluence
3w
Save
Mark Applied
Hide
AI Infrastructure Engineer
Chicago or Hartford or Nashville or Costa Mesa or San Jose or Atlanta or Boston or Cleveland or Columbus or Dallas or Denver or Fort Lauderdale or Grand Rapids or Indianapolis or Los Angeles or Miami or New York or Oakbrook Terrace or Sacramento or San Francisco or South Bend or Tampa or Houston or Austin or Charlotte
$95k-$213k/yr OnsiteFull Time
Crowe
Crowe: Global professional services firm providing audit, tax, and consulting.
Design, operate, and scale production Kubernetes (AKS) platforms; build Terraform and Crossplane infra-as-code; implement GitOps with FluxCD; create Azure DevOps pipelines; define observability (Azure Monitor, Prometheus, Grafana); bachelor’s degree required.
Terraform, Crossplane, FluxCD, GitOps, ADO, Azure DevOps, AKS, Kubernetes, Microsoft Azure, Azure Monitor, Log Analytics, Prometheus, Grafana, OpenFaaS, ARM templates, Octopus Deploy, Key Vault, ACR
1mo
Save
Mark Applied
Hide
Cloud Infrastructure and AI Efficiency Engineer
San Francisco, California, United States
OnsiteFull Time
Apple
AppleNASDAQ: AAPL: Designs and sells consumer electronics, software, and online services.
Engineering role focused on cloud infrastructure and AI efficiency; specific experience and certifications not provided in the description.
1mo
Save
Mark Applied
Hide
Gen AI Infrastructure Engineer
Woodbridge or New York City or Atlanta or Boston or Chicago or Dallas or Delaware or Denver or Garden City or Cayman Islands or Greenwich or Houston or Los Angeles or Miami or Naples or Nevada or Palm Beach or San Diego or San Francisco or Seattle or Stuart or Washington
$160k-$200k/yr HybridFull Time
Bessemer Trust
Bessemer Trust: Wealth management and family office services for affluent clients.
7+ YOE7+ years in DevOps/platform/cloud infrastructure engineering; strong AWS (IAM, CloudFormation, Lambda, API Gateway, VPC, CloudWatch, SSM, Secrets Manager, ECR); AWS CDK/CloudFormation/Terraform; CI/CD (Bitbucket/GitHub Actions, OIDC); container and datastore operations; security fundamentals.
AWS Bedrock, AgentCore, Lambda, API Gateway, AWS CDK, CloudFormation, Terraform, Bitbucket Pipelines, GitHub Actions, OIDC, IAM, VPC, CloudWatch, SSM, Secrets Manager, ECR, Neo4j, Neptune, Redis, Milvus, VectorDB
1mo
Save
Mark Applied
Hide
Senior AI Infrastructure Engineer - DGX Cloud
Santa Clara or California
$152k-$288k/yr HybridFull Time
NVIDIA
NVIDIANASDAQ: NVDA: Designs graphics processing units and artificial intelligence hardware.
5+ YOEBS or equivalent, 5+ years experience, background in infrastructure automation and distributed systems, proficiency with Python/Go/C/C++/Java, Kubernetes/OpenStack, Terraform, Linux, and experience with large-scale cloud platforms.
DGX Cloud, Kubernetes, OpenStack, Python, Go, C/C++, Java, Terraform, Slurm, Linux
3w
Save
Mark Applied
Hide
AI Infrastructure Manager
Bengaluru or San Francisco or Boston or New York City or Austin or Tokyo or London
HybridFull Time
Postman
Postman: Platform for building, testing, and managing software APIs.
Experience leading engineering teams building GenAI or AI infrastructure and distributed systems; strong cloud, accelerator, and performance optimization knowledge; proficiency in Python or Go; architecture and reliability experience.
Python, Go
1mo
Save
Mark Applied
Hide
Senior AI Infrastructure Engineer - DGX Cloud
Santa Clara or California or United States
$152k-$288k/yr OnsiteFull Time
NVIDIA
NVIDIANASDAQ: NVDA: Designs GPU-accelerated computing and artificial intelligence hardware.
5+ YOEBS in CS or related (or equivalent experience), 5+ years experience, expertise in infrastructure automation and distributed systems, proficiency in Python/Go/C/C++/Java, Kubernetes, OpenStack, Terraform, Linux, and capacity/performance management.
Kubernetes, OpenStack, Python, Go, C/C++, Java, Terraform, Slurm, Infrastructure as a Code (IAAC), Linux
3w
Save
Mark Applied
Hide
Infrastructure Engineer
Menlo Park, California, United States
OnsiteFull Time
Shakudo
Shakudo: Develops an operating system for enterprise AI applications.
8+ YOE8+ years engineering experience, 5+ years Kubernetes operation, proficiency in Rust, experience with production infrastructure (physical servers, GPU/DGX clusters), CI/CD, security hardening, observability, and LLM/AI infrastructure.
Kubernetes, Rust, CI/CD, DGX, GPU, LLM, ETL
1mo
Save
Mark Applied
Hide
AI Infrastructure Engineer
San Jose, California, United States
$192k-$250k/yr OnsiteFull Time
NIO
NIONYSE: NIO: Designs and manufactures premium smart electric vehicles and technology
5+ YOE5+ years building and optimizing large-scale LLM/VLM inference systems; strong C/C++ and performance engineering skills; GPU/NPU programming (CUDA), PyTorch/TensorFlow, and BS/MS in CS/CE or related field required.
CUDA, PyTorch, TensorFlow, C/C++, AIOS
1w
Save
Mark Applied
Hide
Senior AI Infrastructure Engineer - Model Training
Mountain View, California, United States
$190k-$260k/yr OnsiteFull Time
Kodiak Robotics
Kodiak RoboticsNASDAQ: KDK: Develops autonomous driving technology for commercial trucking and defense.
2+ YOEDegree in CS or related field,2+ years ML systems experience,expertise in distributed training,high-performance data pipelines,GPU performance and profiling,Python and PyTorch skills.
PyTorch, PyTorch DDP/FSDP, DeepSpeed, Megatron, NCCL, WebDataset, MosaicML Streaming, MDS, Nsight, PyTorch Profiler, Python, C++, CUDA, Triton, NVLink, InfiniBand
3w
Save
Mark Applied
Hide
Lead Infrastructure Engineer-Network Engineer
San Francisco or Seattle
$143k-$185k/yr OnsiteFull Time
JPMorgan Chase
JPMorgan ChaseNYSE: JPM: Global financial services firm providing banking and investment solutions.
5+ YOE5+ years infrastructure engineering experience, formal training/certification, deep cloud and network knowledge, scripting and automation experience, familiarity with security/segmentation and AI-assisted engineering, strong problem-solving and mentoring skills.
Cisco, Juniper, Arista, Juniper Mist, Cisco/Viptela, Fortinet, BGP, OSPF, MPLS, EVPN/VXLAN, SD-WAN, Wi Fi 6E, Wi Fi 7, 5G, Palo Alto, Zscaler, IPsec, TLS, ZTNA, Python, Ansible, Terraform, Git, GitHub Actions, Jenkins, ThousandEyes, Splunk, Grafana, Wireshark
3mo
Save
Mark Applied
Hide
Sr. Cloud AI Infrastructure Engineer
Palo Alto, California, United States
$145k-$273k/yr OnsiteFull Time
Tencent
TencentHong Kong Stock Exchange: 0700: Developing digital services and entertainment for a global audience.
Master’s or Ph.D. in Computer Engineering, Electronic Engineering, Microelectronics, or related field; expertise in GPGPU/AI accelerator architectures; proficient in CUDA and Triton; strong distributed systems knowledge; experience with PyTorch or TensorFlow.
CUDA, Triton, PyTorch, TensorFlow
1mo
Save
Mark Applied
Hide
Forward Deployed Infrastructure Engineer
San Francisco, California, United States
OnsiteFull Time
Sirius Technology: AI-powered retention platform for subscription-based businesses.
1+ YOE1+ year in infrastructure/DevOps/platform engineering or solutions architecture with customer-facing deployment experience; deep experience with AWS, Terraform, container orchestration, and cloud networking (VPC, IAM, DNS); strong operational and stakeholder communication skills.
AWS, Terraform, VPC, IAM, DNS, container orchestration, CRM, AI/ML, LLM
3mo
Save
Mark Applied
Hide
AI Training Infrastructure Engineer – Humanoid Whole Body Control
San Jose, California, United States
$200k-$300k/yr OnsiteFull Time
Figure
Figure: Develops autonomous humanoid robots for commercial and residential tasks.
Proficient in Python and PyTorch; experience with robotics training infra; knowledge of physics engines; RL/IL; dynamics and controls.
Python, PyTorch, NVIDIA PhysX, MuJoCo, Warp, PyBullet, Robotics simulation tools
1mo
Save
Mark Applied
Hide
AI Engineer, Agent Infrastructure
San Francisco, California, United States
OnsiteFull Time
Zed
Zed: AI-native neobank providing premium credit services to young professionals.
Experience shipping production LLM/agent systems; strong backend/infrastructure; familiarity with workflows, observability, and production readiness.
Python, Go, REST APIs, Docker, Kubernetes, LLMs, Observability, Monitoring
2w
Save
Mark Applied
Hide
Senior MLOps & AI Infrastructure Engineer
San Jose, California, United States
$149k-$216k/yr OnsiteFull Time
Altera
Altera: Manufacturer of field-programmable gate arrays and programmable logic devices.
10+ YOEBachelor's degree, 10+ years ML engineering/MLOps experience, strong Python, cloud ML platforms (AWS/GCP/Azure), Docker/Kubernetes, CI/CD, ML frameworks (PyTorch/TensorFlow/JAX), experience with MLflow/W&B and HPC schedulers.
PyTorch, TensorFlow, JAX, Hugging Face, scikit-learn, XGBoost, MLflow, Kubeflow, Airflow, Weights & Biases, DVC, Feast, AWS SageMaker, GCP Vertex AI, Azure ML, Terraform, CloudFormation, Docker, Kubernetes, Slurm, LSF, Python, Bash, Go, SQL, Prometheus, Grafana, ELK Stack, Evidently AI, Arize
2mo
Save
Mark Applied
Hide
Principal Engineer, AI Platform & Infrastructure
San Francisco, California, United States
HybridFull Time
SpreeAI
SpreeAI: AI-powered virtual try-on and sizing software for fashion retailers.
10+ YOE10+ years software engineering/infrastructure; 5+ years ML infrastructure, MLOps, or AI platform engineering; strong Python, PyTorch, Kubernetes, Docker; distributed systems expertise; experience with ML workflow orchestration and production inference systems.
Python, PyTorch, Kubernetes, Docker, cloud infrastructure, GPU workloads