63 reliability engineering manager jobs at 31 companies in Bellevue, WA
1w
Save
Mark Applied
Hide
1w
Engineering Manager, Cloud Network Reliability
Seattle, Washington, United States
OnsiteFull Time
AppleNASDAQ: AAPL: Designs and sells consumer electronics, software, and online services.
Experienced engineering manager required to lead and grow reliability engineers delivering highly available, scalable, resilient, fault-tolerant global network services.
CoreWeaveNASDAQ: CRWV: Cloud platform providing GPU-accelerated infrastructure for AI workloads.
7+ YOE2+ Mgmt7+ years in software or infrastructure engineering with 2+ years leadership; SRE fundamentals, incident management, observability, change management; strong automation and people development skills.
Rippling: Unified platform managing workforce HR, IT, and finance operations
8+ YOE4+ Mgmt8+ years software/infrastructure engineering, 4+ years engineering management, expertise in AWS and Kubernetes, experience with Terraform and Helm, production reliability and incident response experience, strong cross-functional leadership.
Crusoe: Provides energy-efficient cloud infrastructure powered by stranded and renewable energy.
6+ YOE2+ MgmtRequires 6–8+ years in distributed systems or cloud networking engineering, 2–4+ years managing engineering talent, and expertise in SDN, network virtualization, control planes, reliability, and large-scale cloud services.
Senior Software Engineer, Site Reliability Engineering
San Francisco or San Jose or New York City or Seattle or Austin or Washington or California or Massachusetts or New Jersey or Washington or United States
$179k-$273k/yrRemoteFull Time
Thumbtack: Online marketplace connecting homeowners with local service professionals.
5+ YOE5+ years managing infrastructure and systems; extensive AWS and Linux fluency; proficiency in Python, Go, PHP, and JavaScript; experience with distributed systems, observability, and on-call rotations; strong communication and troubleshooting skills.
Anduril Industries: Defense technology building autonomous military hardware and software.
7+ YOE2+ Mgmt7+ years in software/reliability engineering, 2+ years managing software engineers, experience with production systems, AI-enabled software, CI/CD and observability, strong leadership and communication, U.S. Person required.
AmazonNASDAQ: AMZN: Global online retail and cloud computing technology provider.
7+ YOE2+ MgmtBachelor's in Mechanical/Aerospace/Manufacturing Engineering or equivalent;7+ years mechanical engineering experience on complex hardware;2+ years people management;experience with aerospace/high-reliability hardware;US export-control eligibility required.
Livingston or New York City or Sunnyvale or Bellevue or San Francisco
$165k-$242k/yrOnsiteFull Time
CoreWeaveNASDAQ: CRWV: Specialized cloud infrastructure provider for high-performance AI workloads.
5+ YOE3+ MgmtRequires 3+ years managing engineering teams, 5+ years backend software engineering, Go or comparable language, distributed systems, data modeling, production reliability, roadmap planning, and engineer development.
Go, ClickHouse, S3, GCS, Azure, CoreWeave AI Object Storage, RBAC
SpaceX: Designs and launches advanced rockets and satellite internet constellations.
5+ YOE5+ years SRE/DevOps experience with Kubernetes and Linux, proficiency in Bash/Python, experience managing infrastructure and automation, and ability to obtain Top Secret/SCI or DOE Q clearance.
Regional Reliability Engineer II - Equipment/Facilities (Bellevue, WA, US, 98203)
Bellevue, Washington, United States
$110k-$143k/yrOnsiteFull Time
CintasNasdaq: CTAS: Provides managed uniform programs and facility services to businesses.
5+ YOE5+ years leadership in industrial maintenance, equipment installation, or reliability engineering; multi-site change leadership; capital project support; CMMS; AutoCAD; willingness to travel up to 60%.
SnowflakeNYSE: SNOW: Cloud-based platform for data storage, processing, and analytics.
10+ YOE1+ Mgmt10+ years software engineering, 1+ year technical people leadership, systems thinking, experience with Kubernetes-based infrastructure, CI and developer tooling, strong communication and reliability focus.
Sr. Director of Engineering, Managed PostgreSQL and AI-Native Database Platform
Seattle, Washington, United States
$262k-$327k/yrHybridFull Time
DigitalOceanNew York Stock Exchange: DOCN: Simplifies cloud infrastructure for developers, startups, and SMBs.
15+ YOE6+ Mgmt15+ years engineering experience, 6+ years managing managers, deep distributed systems and database expertise, proven production reliability operations, strategic roadmap and executive communication skills.
Tech Lead Cloud Site Reliability Engineer - DCS Cloud
Seattle, Washington, United States
OnsiteFull Time
ByteDance: Developing AI-driven content platforms and mobile applications.
5+ YOEBachelor's in CS or related,5+ years SRE/Linux/DevOps experience,proficient in Go/Python/C++,familiar with public cloud platforms,monitoring,incident response,and strong troubleshooting and communication skills.
CrowdStrikeNASDAQ: CRWD: Provides cloud-native endpoint protection and cybersecurity services.
10+ YOE10+ years building distributed systems, 5+ years developing SaaS microservices, expert programming skills, distributed-systems expertise, architectural leadership, and a Computer Science degree or equivalent experience.
TikTok USDS Joint Venture: Operates and secures TikTok services for U.S. users.
3+ YOEBachelor's in CS or related and 3+ years SRE/systems experience; expertise in Go/Python, Linux, distributed systems, networking, debugging, and collaboration; on-call and incident management experience.
Staff+ Site Reliability Engineer, Safeguards ML Infra
San Francisco or Seattle or New York City
$405k-$485k/yrHybridFull Time
Anthropic: Developing safe and reliable artificial intelligence systems.
8+ YOEProduction change-management experience, high-stakes release and on-call experience, AWS/GCP operations, Python proficiency, and a bachelor's degree or equivalent experience.
Python, Rust, AWS, GCP, AWS Bedrock, GCP Vertex, Claude
Senior System Architect, Infrastructure Reliability
Santa Clara or Westford or Austin or Durham or Redmond
$184k-$357k/yrHybridFull Time
NVIDIANASDAQ: NVDA: Designs GPU-accelerated computing and artificial intelligence hardware.
6+ YOE6+ years systems programming experience, BS/MS/PhD in CS or EE (or equivalent), expertise in CPU/GPU diagnostics, C++ and Python proficiency, experience with RCA, cluster managers (Slurm/LSF/Kubernetes).