95 reliability engineering manager jobs at 45 companies in Marina, CA
1mo
Save
Mark Applied
Hide
1mo
Reliability Engineering Manager - (M5)
Santa Clara, California, United States
$172k-$236k/yrOnsiteFull Time
Applied MaterialsNASDAQ: AMAT: Manufacturers of equipment for semiconductor and display production.
Managing reliability verification and qualification for semiconductor products, leading and developing reliability teams, subject-matter expertise in system-level reliability, familiarity with semiconductor fabrication equipment, and developing AI/Copilot automation.
OracleNYSE: ORCL: Provides cloud infrastructure and enterprise software for global businesses.
5+ YOE3+ Mgmt5+ years network reliability engineering, 3+ years engineering/operations management, strong cloud networking and distributed systems expertise, proven people leadership, excellent communication and organizational skills.
Site Reliability Engineering Manager, Storage - Apple Services Engineering
Cupertino, California, United States
OnsiteFull Time
AppleNASDAQ: AAPL: Designs and sells consumer electronics, software, and online services.
Engineering manager for distributed storage systems and large-scale storage infrastructure; experience in distributed systems and site reliability engineering.
NVIDIANASDAQ: NVDA: Designs GPU-accelerated computing and artificial intelligence hardware.
18+ YOE10+ Mgmt18+ years in board and system reliability, 5+ years on data center equipment, 10+ years leading reliability management; deep reliability and physics-of-failure expertise; statistics and reliability modeling skills; bachelor’s in engineering or related (graduate preferred).
NVIDIANASDAQ: NVDA: Designs graphics processing units and artificial intelligence hardware.
18+ YOE10+ Mgmt18+ years in board/system reliability with 5+ years on data center equipment and 10+ years leading reliability; deep reliability, testing, modeling, statistics, and physics-of-failure expertise.
DoorDashNASDAQ: DASH: On-demand delivery platform connecting consumers with local merchants.
5+ YOE5+ Mgmt5+ years leading engineering teams and 5+ years in infrastructure/platform/backend roles; strong platform mindset, SRE experience (SLOs/error budgets), AWS and cloud fundamentals, influence and hiring experience, experience with incident/incident response processes.
8+ YOE2+ MgmtBachelor's degree or equivalent, 8+ years in software/systems/SRE, 5+ years building large-scale infrastructure, 2+ years people management; Master's degree preferred.
5+ YOE3+ Mgmt5+ years software engineering, 3+ years managing engineering teams; hands-on Kubernetes, Terraform, AWS; experience with production reliability, SLOs, on-call and incident response.
NetflixNASDAQ: NFLX: Global video streaming and media production service.
Lead a team of distributed systems engineers; build and operate large-scale payments services; collaborate with product, regional leads, and stakeholders; drive reliability, scalability, and innovation.
UberNYSE: UBER: A technology platform for transportation, delivery, and freight.
10+ YOE2+ MgmtProven engineering leadership with strong backend system design skills, 10+ years engineering experience, 2+ years managing engineering teams, hiring and developing talent, and ownership of reliability and operational excellence.
Senior Software Engineer, Site Reliability Engineering
San Francisco or San Jose or New York City or Seattle or Austin or Washington or California or Massachusetts or New Jersey or Washington or United States
$179k-$273k/yrRemoteFull Time
Thumbtack: Online marketplace connecting homeowners with local service professionals.
5+ YOE5+ years managing infrastructure and systems; extensive AWS and Linux fluency; proficiency in Python, Go, PHP, and JavaScript; experience with distributed systems, observability, and on-call rotations; strong communication and troubleshooting skills.
NetflixNASDAQ: NFLX: Provider of global streaming entertainment and video content.
3+ Mgmt3+ years managing distributed engineering teams, platform mindset, experience with reliable 24x7 services, strong communication, hiring/coaching, and ability to drive adoption of data platforms and governance.
graph-based data modeling, knowledge graphs, RDF, ontologies
Santa Clara or St. Louis or Bangalore or London or Paris or Melbourne or Taipei or Tokyo
OnsiteFull Time
NetskopeNASDAQ: NTSK: Cloud-native cybersecurity and data protection platform for enterprises.
3+ YOEBachelor's in CS/Engineering or equivalent; 3+ years building/managing complex systems (including 1-2 years SRE); experience with cloud services, microservices, availability/performance optimization, debugging, and strong communication.
Contract Lead, Site Reliability Engineering — AI Accelerator Infrastructure
Santa Clara, California, United States
$195k-$285k/yrHybridContract, Full Time
d-Matrix: Develops high-performance semiconductor chips for generative AI inference.
15+ YOE5+ MgmtBachelor's in CS/EE,15+ years SRE/infrastructure engineering,5+ years leading SRE teams,deep Linux,Terraform,Ansible,Kubernetes,Prometheus/Grafana/Datadog,Python or Go,cloud (AWS/Azure/GCP).
AbbottNYSE: ABT: Manufactures medical devices, diagnostics, and nutritional health products.
Ensure reliability, scalability, and performance of a medical-device remote monitoring platform; expertise in cloud (Azure), Kubernetes, observability, automation, and incident management; bachelor's in a technical discipline.
Python, Go, Bash, PowerShell, Microsoft Azure, Azure Kubernetes Service (AKS), Azure Monitor, Azure DevOps, Azure Policy, Kubernetes, Docker, Prometheus, Grafana, ELK/EFK, Datadog, Linux
Altera: Manufacturer of field-programmable gate arrays and programmable logic devices.
15+ YOE5+ MgmtLead global hardware engineering; 15+ years in semiconductor hardware; manage US/Asia teams; expertise in packaging, board hardware, SI/PI, and reliability.