570 ai infrastructure engineer jobs at 237 companies in Berkeley, CA
🚀PromotedHiringCafe
ML Engineer - Inference & Model Deployment
Cupertino, CA, US
$250k-$310k/yrOn-SiteFull Time
HiringCafe: Building a 100× better job search engine to take on Indeed and LinkedIn.
Turn powerful AI and ML models into fast, reliable production systems. Own inference latency, throughput, model-serving architecture, multi-GPU systems, and production deployment for millions of users.
HiringCafe: Building a 100× better job search engine to take on Indeed and LinkedIn.
Build the ML and AI search behind HiringCafe — ranking, recommenders, retrieval, and LLM agents that surface jobs people would never find on their own.
HiringCafe: Building a 100× better job search engine to take on Indeed and LinkedIn.
Own the crawlers, pipelines, and infrastructure powering a real-time job search engine. Strong Node.js and Python fundamentals; bonus points for security and reverse-engineering chops.
WiproNYSE: WIT: Global technology services and consulting for digital transformation.
3+ YOEDesign and build hybrid AI infrastructure integrating compute, GPU nodes, software-defined storage, Kubernetes, IaC, and observability; 3+ years experience in AI infrastructure or related roles.
Together AI: Cloud platform for training and deploying artificial intelligence models.
5+ YOE5+ years in AI infrastructure or related roles; BS in CS or equivalent; knowledge of Ansible, Terraform, Kubernetes; programming/scripting; monitoring/observability; cloud services; collaborative work
AMDNASDAQ: AMD: Designs and manufactures computer processors and graphics technology.
5+ YOE5+ years in DevOps/platform/infrastructure engineering; deep Kubernetes experience; experience building developer-facing platforms, Helm and GitOps workflows, storage/networking for GPU workloads; Terraform, monitoring, and ML framework exposure preferred.
Chicago or Hartford or Nashville or Costa Mesa or San Jose or Atlanta or Boston or Cleveland or Columbus or Dallas or Denver or Fort Lauderdale or Grand Rapids or Indianapolis or Los Angeles or Miami or New York or Oakbrook Terrace or Sacramento or San Francisco or South Bend or Tampa or Houston or Austin or Charlotte
$95k-$213k/yrOnsiteFull Time
Crowe: Global professional services firm providing audit, tax, and consulting.
Design, operate, and scale production Kubernetes (AKS) platforms; build Terraform and Crossplane infra-as-code; implement GitOps with FluxCD; create Azure DevOps pipelines; define observability (Azure Monitor, Prometheus, Grafana); bachelor’s degree required.
Woodbridge or New York City or Atlanta or Boston or Chicago or Dallas or Delaware or Denver or Garden City or Cayman Islands or Greenwich or Houston or Los Angeles or Miami or Naples or Nevada or Palm Beach or San Diego or San Francisco or Seattle or Stuart or Washington
$160k-$200k/yrHybridFull Time
Bessemer Trust: Wealth management and family office services for affluent clients.
7+ YOE7+ years in DevOps/platform/cloud infrastructure engineering; strong AWS (IAM, CloudFormation, Lambda, API Gateway, VPC, CloudWatch, SSM, Secrets Manager, ECR); AWS CDK/CloudFormation/Terraform; CI/CD (Bitbucket/GitHub Actions, OIDC); container and datastore operations; security fundamentals.
NVIDIANASDAQ: NVDA: Designs graphics processing units and artificial intelligence hardware.
5+ YOEBS or equivalent, 5+ years experience, background in infrastructure automation and distributed systems, proficiency with Python/Go/C/C++/Java, Kubernetes/OpenStack, Terraform, Linux, and experience with large-scale cloud platforms.
Bengaluru or San Francisco or Boston or New York City or Austin or Tokyo or London
HybridFull Time
Postman: Platform for building, testing, and managing software APIs.
Experience leading engineering teams building GenAI or AI infrastructure and distributed systems; strong cloud, accelerator, and performance optimization knowledge; proficiency in Python or Go; architecture and reliability experience.
NVIDIANASDAQ: NVDA: Designs GPU-accelerated computing and artificial intelligence hardware.
5+ YOEBS in CS or related (or equivalent experience), 5+ years experience, expertise in infrastructure automation and distributed systems, proficiency in Python/Go/C/C++/Java, Kubernetes, OpenStack, Terraform, Linux, and capacity/performance management.
Kubernetes, OpenStack, Python, Go, C/C++, Java, Terraform, Slurm, Infrastructure as a Code (IAAC), Linux
Shakudo: Develops an operating system for enterprise AI applications.
8+ YOE8+ years engineering experience, 5+ years Kubernetes operation, proficiency in Rust, experience with production infrastructure (physical servers, GPU/DGX clusters), CI/CD, security hardening, observability, and LLM/AI infrastructure.
NIONYSE: NIO: Designs and manufactures premium smart electric vehicles and technology
5+ YOE5+ years building and optimizing large-scale LLM/VLM inference systems; strong C/C++ and performance engineering skills; GPU/NPU programming (CUDA), PyTorch/TensorFlow, and BS/MS in CS/CE or related field required.
Senior AI Infrastructure Engineer - Model Training
Mountain View, California, United States
$190k-$260k/yrOnsiteFull Time
Kodiak RoboticsNASDAQ: KDK: Develops autonomous driving technology for commercial trucking and defense.
2+ YOEDegree in CS or related field,2+ years ML systems experience,expertise in distributed training,high-performance data pipelines,GPU performance and profiling,Python and PyTorch skills.
JPMorgan ChaseNYSE: JPM: Global financial services firm providing banking and investment solutions.
5+ YOE5+ years infrastructure engineering experience, formal training/certification, deep cloud and network knowledge, scripting and automation experience, familiarity with security/segmentation and AI-assisted engineering, strong problem-solving and mentoring skills.
Cisco, Juniper, Arista, Juniper Mist, Cisco/Viptela, Fortinet, BGP, OSPF, MPLS, EVPN/VXLAN, SD-WAN, Wi Fi 6E, Wi Fi 7, 5G, Palo Alto, Zscaler, IPsec, TLS, ZTNA, Python, Ansible, Terraform, Git, GitHub Actions, Jenkins, ThousandEyes, Splunk, Grafana, Wireshark
TencentHong Kong Stock Exchange: 0700: Developing digital services and entertainment for a global audience.
Master’s or Ph.D. in Computer Engineering, Electronic Engineering, Microelectronics, or related field; expertise in GPGPU/AI accelerator architectures; proficient in CUDA and Triton; strong distributed systems knowledge; experience with PyTorch or TensorFlow.
Sirius Technology: AI-powered retention platform for subscription-based businesses.
1+ YOE1+ year in infrastructure/DevOps/platform engineering or solutions architecture with customer-facing deployment experience; deep experience with AWS, Terraform, container orchestration, and cloud networking (VPC, IAM, DNS); strong operational and stakeholder communication skills.
Altera: Manufacturer of field-programmable gate arrays and programmable logic devices.
10+ YOEBachelor's degree, 10+ years ML engineering/MLOps experience, strong Python, cloud ML platforms (AWS/GCP/Azure), Docker/Kubernetes, CI/CD, ML frameworks (PyTorch/TensorFlow/JAX), experience with MLflow/W&B and HPC schedulers.
SpreeAI: AI-powered virtual try-on and sizing software for fashion retailers.
10+ YOE10+ years software engineering/infrastructure; 5+ years ML infrastructure, MLOps, or AI platform engineering; strong Python, PyTorch, Kubernetes, Docker; distributed systems expertise; experience with ML workflow orchestration and production inference systems.