225 reliability engineering manager jobs at 145 companies in El Cerrito, CA
1mo
Save
Mark Applied
Hide
1mo
Senior Director, Reliability Engineering
Santa Clara, California, United States
$332k-$500k/yrOnsiteFull Time
NVIDIANASDAQ: NVDA: Designs GPU-accelerated computing and artificial intelligence hardware.
18+ YOE10+ Mgmt18+ years in board and system reliability, 5+ years on data center equipment, 10+ years leading reliability management; deep reliability and physics-of-failure expertise; statistics and reliability modeling skills; bachelor’s in engineering or related (graduate preferred).
OktaNASDAQ: OKTA: Provide secure identity management and authentication for enterprises.
3+ Mgmt3+ years technical leadership experience; experience with cloud-native architectures, Kubernetes, Terraform, CI/CD, observability platforms; strong software development and automation background; US Person status required.
Amazon Web Services (AWS), Kubernetes, Terraform, Grafana, Splunk, APM, CI/CD
NVIDIANASDAQ: NVDA: Designs graphics processing units and artificial intelligence hardware.
18+ YOE10+ Mgmt18+ years in board/system reliability with 5+ years on data center equipment and 10+ years leading reliability; deep reliability, testing, modeling, statistics, and physics-of-failure expertise.
Wells FargoNYSE: WFC: Global provider of banking, investment, and mortgage financial services.
7+ YOE3+ MgmtRequires 7+ years in systems engineering or technology architecture, 3+ years of management, 5+ years leading engineering or SRE teams, and experience with customer-facing platforms, incident management, SRE, DevOps, cloud, and production operations.
Splunk, Grafana, AppDynamics, Dynatrace, OpenTelemetry, Prometheus, Kubernetes, OpenShift, AWS, Microsoft Azure, Google Cloud Platform, Infrastructure as Code (IaC), CI/CD, AIOps, ITIL
CiscoNASDAQ: CSCO: Develops and sells networking hardware and cybersecurity software.
8+ YOEBachelor's in Engineering with 12+ years or Master's with 8+ years; deep hardware reliability and PCBA knowledge; expertise in RAS, risk management, data-driven reliability, and executive influence; proven leadership and mentoring skills.
Principal Tech Lead Manager - Data Platform & Reliability Engineering
Mountain View, California, United States
$215k-$275k/yrOnsiteFull Time
ID.me: Provides secure digital identity verification and authentication services.
5+ YOE3+ Mgmt8+ years engineering experience with 3+ years managing teams,5+ years in data/platform/SRE; bachelor\u0002s or equivalent; deep PostgreSQL and data reliability expertise; strong communication and cloud/IaC experience.
Forge GlobalNYSE: FRGE: Marketplace for trading private shares and pre-IPO stock.
8+ YOE5+ Mgmt8+ years software engineering experience with infrastructure/platform focus, 5+ years people leadership, cloud and infrastructure as code experience, observability and incident response expertise, Bachelor's in CS or equivalent, strong communication.
Site Reliability Engineering (SRE) Manager, Apple Maps
Cupertino, California, United States
OnsiteFull Time
AppleNASDAQ: AAPL: Designs and sells consumer electronics, software, and online services.
Build, manage, and deliver highly available, automated infrastructure for Apple Maps at global scale; focus on reliability, scalability, and operational excellence.
Berlin or Toronto or San Francisco or New York City or London or Paris or Montreal or Seoul or Germany
HybridFull Time
Cohere: Provides enterprise-grade large language models and AI software platforms.
Proven experience managing engineering teams, building and shipping production AI-first full-stack applications, partnering with product and ML teams, and delivering secure, reliable systems.
LiteLLM: Open-source AI gateway for standardizing LLM API access.
Proven engineering manager with experience shipping through teams, metrics-driven reliability/process design, incident/RCAs, hiring and scaling, customer-facing escalation handling, and familiarity with Rust migrations and startup/open-source environments.
Checkr: AI-powered platform for background checks and identity verification.
8+ YOE4+ MgmtRequires 4+ years managing engineering teams, 8+ years as a software engineer, service architecture and distributed systems expertise, integration reliability, incident management, regulated products, and strong stakeholder communication.
Rec Technologies: Modern software platform for parks and recreation departments.
5+ YOE4+ Mgmt5+ years as a software engineer with 4+ years managing engineering teams, AI fluency, experience with platform reliability or payments, strong coding and hiring skills, and ability to coach engineers and shape roadmap.
10+ YOE3+ Mgmt10+ years in hardware test/qualification or reliability engineering with 3+ years leading teams, BS in engineering or CS, hands-on spaceflight or environmental test experience, reliability analysis knowledge, and strong communication skills.
Render: Cloud platform for building, deploying, and scaling AI-native applications.
8+ YOE4+ Mgmt8+ years building infrastructure or platform products for developers, 4+ years managing engineers, strong system design, reliability, CI/CD, observability, configuration management, and developer experience.
Pathstream: Online workforce development and professional certificate programs.
4+ YOE2+ Mgmt4+ years software engineering experience, 2+ years leading engineers, proficiency in modern web stacks and cloud (Ruby on Rails, React, JS/TS, Python, Docker, PostgreSQL, AWS), experience with production reliability, security, and AI-enabled development tools.
Ruby on Rails, React, JavaScript, TypeScript, Python, Docker, PostgreSQL, AWS, Claude Code
Fivetran: Automates data movement into cloud data warehouses.
Experience managing software engineering teams, reviewing designs and code, delivering complex cloud projects, owning production services, driving reliability, planning with product partners, and using AI development tools.
Hightouch: Syncs customer data from warehouses to business and marketing tools.
Lead an engineering team for an Agentic Marketing Platform: roadmap and ship features, hire and grow engineers, improve reliability and execution, and provide strong technical and people leadership.