113 cloud reliability engineer jobs at 41 companies in Washington

1mo
Save
Mark Applied
Hide
Tech Lead Cloud Site Reliability Engineer - DCS Cloud
Seattle, Washington, United States
OnsiteFull Time
ByteDance
ByteDance: Global technology specializing in AI-powered content platforms.
5+ YOEBachelor's in CS or related,5+ years SRE/Linux/DevOps experience,proficient in Go/Python/C++,familiar with public cloud platforms,monitoring,incident response,and strong troubleshooting and communication skills.
Linux, Go, Python, C++, OCI, AWS, Azure, GCP, KVM/QEMU, Docker, Kubernetes, containerd, cgroups, namespaces, CUDA, MIG
3w
Save
Mark Applied
Hide
Systems Reliability Engineer
Overland Park or Atlanta or Frisco or Bellevue
$84k-$151k/yr OnsiteFull Time
T-Mobile
T-MobileNASDAQ: TMUS: The Un-carrier providing wireless and home internet services.
2+ YOEBachelor's degree required; 2–4+ years preferred. Requires DevOps, cloud, automation, monitoring, scripting, APIs, cybersecurity, and reliability engineering experience, plus U.S. work authorization.
C, C#, Java, Perl, Python, Go, Jenkins, CloudBees, Ansible, Chef, Puppet, Docker, Kubernetes, AppDynamics, Splunk, Microsoft Graph API, REST API, Microsoft Power Apps, Microsoft Power Automate, Microsoft Entra, SailPoint, ServiceNow, Azure, Microsoft Azure DevOps Pipelines, Linux, Shell
2mo
Save
Mark Applied
Hide
Site Reliability Engineer (SRE) - Hybrid Cloud Storage
Seattle, Washington, United States
$140k-$210k/yr HybridFull Time
Qumulo
Qumulo: The world's most advanced file system – any data, any location, total control.
3+ YOE3+ years building/operating automated testing for complex software; strong C and Python skills; experience with on‑prem and cloud (AWS/GCP/Azure); Linux/Ubuntu fluency; familiarity with Kubernetes, Ansible, Terraform, and observability tooling.
Python, C, Jenkins, Argo, OpenMetrics, Grafana, InfluxDB, Prometheus, AWS, GCP, Azure, Linux, Ubuntu, Ansible, Terraform, Kubernetes, NFS, SMB, S3
4d
Save
Mark Applied
Hide
Service Reliability Engineer
Liberty Lake, Washington, United States
$80k-$110k/yr HybridFull Time
OpenEye
OpenEye: Private commercial cloud video surveillance serving businesses with AI-driven analytics, business intelligence, and loss-prevention tools.
1+ YOERequires 1–5 years of related experience, cloud, CI/CD, infrastructure automation, monitoring, scripting or development experience, TCP/IP knowledge, Agile familiarity, and strong problem-solving and communication skills.
AWS, Coralogix, TypeScript, MySQL, CrateDB, Git, Java, JavaScript, C#, C++, Datadog, Prometheus, Grafana, Jira
1mo
Save
Mark Applied
Hide
Senior Site Reliability Engineer - Core Cloud Platform
San Francisco or San Jose or Bellevue
$240k-$356k/yr HybridFull Time
Lambda
Lambda: AI infrastructure building GPU cloud services and supercomputers for researchers, enterprises, and hyperscalers.
7+ YOE7+ years SRE or production infrastructure experience, deep Kubernetes and Terraform knowledge, experience with observability and SLOs, proficiency in Go or Python, on-call and incident leadership experience.
Kubernetes, Terraform, Argo CD, Flux, Helm, Kustomize, OpenTelemetry, Prometheus, Grafana, Datadog, Go, Python, etcd, GitOps
2mo
Save
Mark Applied
Hide
Sr. Site Reliability Engineer
Bellevue, Washington, United States
$120k-$150k/yr HybridFull Time
Practice by Numbers
Practice by Numbers: Dental software providing an all-in-one operations, analytics, communications, payments, and marketing platform for dental practices.
6+ YOEEngineering degree (BS/MS) required, 6+ years software/SRE experience, production-quality programming in Go/Python/Java/TypeScript, cloud experience (AWS preferred), on-call and incident leadership experience, SLO/SLI/observability skills.
AWS, Docker, Kubernetes, Terraform, Prometheus, Grafana, OpenTelemetry, Go, Python, TypeScript, GitHub Actions
1mo
Save
Mark Applied
Hide
Senior Site Reliability Engineer I
Boston or Seattle or Atlanta
$134k-$215k/yr HybridFull Time
Axon
AxonNASDAQ: AXON: Develops public safety technologies, devices, and cloud software.
7+ YOEBachelor's in CS/Engineering, 7+ years software engineering experience, expertise in distributed systems, Kubernetes, cloud (Azure/AWS/GCP), observability, Kafka, Terraform/Pulumi, and experience with agentic AI/LLM tooling preferred.
Kubernetes, Terraform, Pulumi, Kafka, Grafana, Datadog, New Relic, MySQL, Cassandra, PostgreSQL, Azure, AWS, GCP
3w
Save
Mark Applied
Hide
Systems Reliability Engineer
Overland Park or Atlanta or Frisco or Bellevue
$84k-$151k/yr OnsiteFull Time
T-Mobile
T-MobileNASDAQ: TMUS: The Un-carrier providing wireless and home internet services.
2+ YOEBachelor's degree required; 2–4+ years preferred. Requires DevOps, cloud, automation, monitoring, Python, APIs, Power Platform, identity governance, and reliability engineering experience.
C, C#, Java, Perl, Python, Go, Shell, Jenkins, CloudBees, Ansible, Chef, Puppet, Docker, Kubernetes, AppDynamics, Splunk, Microsoft Graph API, REST API, Microsoft Power Apps, Microsoft Power Automate, Microsoft Entra, SailPoint, ServiceNow, Azure, Microsoft Azure DevOps Pipelines, Microsoft Power Platform, Linux, VMs
1mo
Save
Mark Applied
Hide
Sr. Site Reliability Engineer (Starlink)
Hawthorne or Palo Alto or Redmond
$165k-$270k/yr OnsiteFull Time
SpaceX
SpaceXNasdaq: SPCX: Designing, manufacturing, and launching advanced rockets and spacecraft.
5+ YOEBachelor's in CS/engineering/math with 5 years software experience or 7+ years SRE/DevOps experience; Linux experience required; Kubernetes, Kafka, cloud-native tooling, and programming in Python/Go/Java/C#/Scala preferred.
Linux, Kubernetes, Istio, Apache Kafka, Apache Spark, HBase, HDFS, Apache Flink, Python, C#, Java, Scala, Go
1mo
Save
Mark Applied
Hide
Senior Site Reliability Engineer, Production Engineer - ThousandEyes
San Francisco or Seattle or Austin or New York City
$165k-$241k/yr HybridFull Time
ThousandEyes
ThousandEyes: ThousandEyes is a Cis-owned digital experience assurance platform helping organizations monitor networks, applications, and cloud services.
5+ YOE5+ years experience; proficiency in Python or Go; expertise with Kubernetes, cloud (AWS), Unix/Linux; strong SRE principles, incident response, and security-minded engineering.
Python, Go, Kubernetes, Service Mesh, Prometheus, OpenTelemetry, ArgoCD, CNCF, AWS, Unix, Linux
1mo
Save
Mark Applied
Hide
Staff Site Reliability Engineer (FedRAMP)
Bellevue or Chicago or New York City or San Francisco or Washington
$174k-$267k/yr HybridFull Time
Okta
OktaNASDAQ: OKTA: Identity management and access control software provider.
8+ YOE8+ years operations experience in cloud and Linux, strong networking and web server knowledge, proficiency with Terraform/Chef and scripting (Bash, Python, Go), experience with automation tools and on-call duty.
AWS, Terraform, Chef, Ansible, Puppet, Apache httpd, nginx, Apache Tomcat, Bash, Python, Golang, git, gdb, strace, ltrace, tcpdump, Wireshark, Docker, Kubernetes
4d
Save
Mark Applied
Hide
Alibaba Cloud-Site Reliability Engineer-Bellevue
Bellevue, Washington, United States
$145k-$238k/yr OnsiteFull Time
Alibaba Cloud
Alibaba CloudNYSE, HKEX: BABA, 9988: Global cloud computing and data intelligence service provider.
5+ YOEBachelor's degree in computer science or related field, 5+ years developing or operating large-scale distributed systems, expert Linux administration, Kubernetes, Python, Shell, and big data architecture expertise.
Linux, Kubernetes, Python, Shell, RAG
1mo
Save
Mark Applied
Hide
Senior Engineer 2 - Site Reliability Engineering (Hybrid, Seattle)
Seattle, Washington, United States
$166k-$258k/yr HybridFull Time
Nordstrom
NordstromNew York Stock Exchange: JWN: Fashion specialty retailer offering apparel, footwear, and accessories.
10+ YOEBachelor's in CS/Engineering or equivalent,10+ years software engineering experience in SRE/infrastructure,proficiency with Kubernetes,cloud providers,networking,strong problem-solving and communication skills.
Kubernetes, Java, Go, Python, AWS, GCP, Azure
2mo
Save
Mark Applied
Hide
Principal Site Reliability Engineer - CTJ - Secret
Redmond, Washington, United States
$143k-$275k/yr OnsiteFull Time
Microsoft
MicrosoftNASDAQ: MSFT: Multinational technology providing software, cloud, and AI solutions.
2+ YOEDegree in CS/IT (or equivalent experience) with minimum 2+ years technical experience (Doctorate path) and ability to obtain required background investigations (T3/CJIS) for government cloud environments; SRE, incident response, and cloud systems experience.
3mo
Save
Mark Applied
Hide
Senior Principal Network Reliability Engineer - Network Region Build (NRB)
Seattle or United States
$126k-$264k/yr OnsiteFull Time
Oracle Corporation
Oracle CorporationNYSE: ORCL: Cloud infrastructure and enterprise software solutions provider.
6+ YOEExpertise in hyperscale cloud networking, routing and transport protocols, reliability engineering, automation, and cross-organizational technical leadership; 6+ years experience preferred; English required.
BGP, OSPF, IS-IS, MPLS, EVPN, VXLAN, IPv4, IPv6, DNS, DHCP, AI
2w
Save
Mark Applied
Hide
Vice President, Reliability
Palo Alto or Vancouver
$320k-$350k/yr OnsiteFull Time
HP Inc.
HP Inc.NYSE: HPQ: Global leader in personal computing, printing, and technology solutions.
15+ YOE8+ MgmtRequires a bachelor's or master's degree, 15+ years in software, infrastructure, platform, or reliability engineering, and 8+ years leading engineering organizations. VP-level experience, cloud-native expertise, and global-scale operations required.
AI, site reliability engineering (SRE)
1w
Save
Mark Applied
Hide
Sr. Staff Engineer Software, Infrastructure Reliability (Chronosphere)
San Francisco or Denver or Austin or Jacksonville or Bridgeport or Seattle or Boston or New York City
$126k-$205k/yr RemoteFull Time
Palo Alto Networks
Palo Alto NetworksNASDAQ: PANW: Global cybersecurity platform providing network, cloud, and AI-driven security solutions.
8+ YOERequires 8+ years of relevant experience, backend programming proficiency, cloud-native and distributed systems expertise, Linux and networking knowledge, debugging skills, and experience with AWS or GCP and Kubernetes.
Go, Java, Python, Rust, AWS, GCP, Kubernetes, Linux, Terraform, Cursor, Claude
1mo
Save
Mark Applied
Hide
Alibaba-Site Reliability Engineer-Bellevue
Bellevue, Washington, United States
$133k-$220k/yr OnsiteFull Time
Alibaba Group
Alibaba GroupNYSE, HKEX: BABA, 9988: Global technology focused on AI, cloud, and consumption.
3+ YOERequires 3+ years in SRE, DevOps, or backend development; proficiency in Python, Go, Java, or C++; Linux, networking, databases, cloud-native systems, incident response, and Chinese-English fluency.
Python, Go, Java, C++, Linux, TCP, HTTP, Prometheus, Istio, Calico, Kubernetes, Alibaba Cloud
3w
Save
Mark Applied
Hide
Engineering Manager, Cloud Network Reliability
Seattle, Washington, United States
OnsiteFull Time
Apple
AppleNASDAQ: AAPL: Designing and manufacturing consumer electronics, software, and digital services.
Experienced engineering manager required to lead and grow reliability engineers delivering highly available, scalable, resilient, fault-tolerant global network services.
2mo
Save
Mark Applied
Hide
Ground System Site Reliability Engineer II
Washington, United States
$135k-$189k/yr OnsiteFull Time
Blue Origin
Blue Origin: Developing reusable space vehicles and infrastructure.
3+ YOEBS or higher in a technical field or equivalent experience, 3+ years software development/SRE experience, proficiency with CI/CD, IaC (Terraform/Pulumi), cloud (AWS/Azure/GCP), Kubernetes, Linux, and at least one language (Python, Go, Rust, Java); U.S. work authorization required.
Terraform, Pulumi, AWS, Azure, GCP, Kubernetes, Istio, Cilium, Linkerd, Linux, Python, Go, Rust, Java, REST, Kafka, gRPC, CI/CD, Software-in-the-Loop, Hardware-in-the-Loop

Explore Jobs

Expand Your Job Search