998 cloud reliability engineer jobs at 502 companies in United States

1mo
Save
Mark Applied
Hide
Cloud Reliability Engineer
Englewood Cliffs or New York City
$135k-$165k/yr HybridFull Time
Versant Media Group
Versant Media GroupNasdaq: VSNT: Media and entertainment managing diverse content and digital brands.
3+ YOEBachelor's degree or equivalent experience; 3–7 years in SRE/Cloud/DevOps roles; strong AWS experience (enterprise scale), Terraform and CloudFormation, scripting (Python/PowerShell/Bash), CI/CD, monitoring/observability, incident management.
AWS, AWS Organizations, Control Tower, Identity Center, Terraform, CloudFormation, Python, PowerShell, Bash, CI/CD
2mo
Save
Mark Applied
Hide
Staff Cloud Reliability Engineer
Irvine or Los Angeles
$180k-$200k/yr OnsiteFull Time
Viant Technology
Viant TechnologyNasdaq Global Select Market: DSP: Public AI-powered advertising platform helping advertisers buy and measure connected-TV and open-internet campaigns programmatically.
8+ YOE8+ years in DevOps/SRE, 3+ years Linux, cloud (AWS/Google), serverless (AWS Lambda/Google Cloud Functions), Docker/Kubernetes, Terraform, CI/CD (GitHub Actions), Python or Go, SQL/BigQuery; participate in on-call rotation.
Linux, AWS, Google, AWS Lambda, Google Cloud Functions, Docker, Kubernetes, Terraform, GitHub Actions, Python, GoLang, SQL, Google BigQuery
1mo
Save
Mark Applied
Hide
Azure Cloud Reliability Engineer III
Salisbury, North Carolina, United States
$125k-$188k/yr HybridFull Time
Ahold Delhaize USA
Ahold Delhaize USAEuronext Amsterdam: AD: The U.S. service division for a global grocery retailer.
5+ YOERequires 5+ years in cloud engineering or platform operations, Azure production experience, infrastructure-as-code skills, automation, networking, IAM, and a bachelor's degree or equivalent experience.
Azure IaaS, Azure PaaS, Azure Functions, Azure Automation, Azure Monitor, Log Analytics, ADO, ARM, Terraform, Ansible, PowerShell, Python, AzCLI, GitHub, containers, CI/CD, Azure AD, PIM, Conditional Access, MFA, Azure AD Connect, Microsoft Defender, Key Vault, Azure Virtual Network, VWAN, ExpressRoute, Load Balancer, Traffic Manager, CDN, Azure DNS, BGP, Azure Virtual Machines, Kubernetes, OpenShift, Azure Storage Account, Disk, Snapshot, Backup, Site Recovery, File Sync, Data Lake
3w
Save
Mark Applied
Hide
Principal Engineer, Cloud Site Reliability Engineering
Santa Clara, California, United States
$272k-$431k/yr OnsiteFull Time
NVIDIA
NVIDIANASDAQ: NVDA: Computing platform for AI and accelerated graphics.
15+ YOEBS/MS in engineering or computer science, 15+ years of systems software development including 1+ year in AI, cloud infrastructure experience, and strong Java, Python, Shell, distributed systems, and database skills.
Java, Python, Shell, REST APIs, MySQL, Cassandra, MongoDB, Elasticsearch, Docker, Virtual Machines, OpenStack, Kubernetes, Chef, Puppet, Hadoop, Ceph, SwiftStack, LXC, Git, Perforce, JFrog, Kafka, Windows, Linux, Android
2mo
Save
Mark Applied
Hide
Senior Engineer, Edge-Cloud Reliability
Reston, Virginia, United States
$191k-$253k/yr OnsiteFull Time
Anduril Industries
Anduril Industries: Defense technology developing AI-powered autonomous military systems.
Active U.S. Top Secret/SCI clearance, strong software background (Go/Rust/Python), distributed-systems fundamentals, cloud architecture depth (GCP preferred), edge/on‑prem comfort, and willingness to travel.
Go, Rust, Python, GCP, Keycloak, LDAP, NixOS, VPC, DNS, Identity Provider (IdP)
1mo
Save
Mark Applied
Hide
Cloud Site Reliability Engineer - DCS Cloud
San Jose, California, United States
OnsiteFull Time
ByteDance
ByteDance: Global technology specializing in AI-powered content platforms.
2+ YOEBachelor's degree in CS or related,2+ years in Linux operations/SRE/DevOps,programming in Go/Python/C++,cloud and reliability practices experience,strong troubleshooting and communication skills.
Go, Python, C++, Linux, OCI, AWS, Azure, GCP, KVM, QEMU, Docker, Kubernetes, containerd, cgroups, namespaces, CUDA, MIG
1mo
Save
Mark Applied
Hide
Staff Network Reliability Engineer (Cloud Operations)
Mountain View, California, United States
HybridFull Time
Skylo Technologies
Skylo Technologies: Private telecommunications providing satellite connectivity for smartphones, vehicles, and IoT devices where cellular networks are unavailable.
8+ YOE8+ years cloud/infrastructure/SRE experience with Kubernetes, hybrid cloud operations, observability, database and storage reliability, GitOps, and on-call ownership in 24x7 environments.
Kubernetes, GKE, GCP, kubectl, Pub/Sub, Cloud SQL, Prometheus, VictoriaMetrics, Grafana, OpenTelemetry, PostgreSQL, Redis, ArgoCD, Helm, Terraform, Ansible, Ceph, Rook, Harvester, KubeVirt, KVM, Loki, ELK, Flux CD, Go, Python, BGP, VXLAN, EVPN
2mo
Save
Mark Applied
Hide
Lead Cloud Engineer, Google Cloud
United States
$177k-$208k/yr RemoteFull Time
Egen
Egen: Technology services helping organizations accelerate value with cloud, data, platforms, and AI.
10+ YOE10+ years cloud engineering experience; expert in Google Cloud; strong Terraform, Kubernetes, CI/CD, cloud security, hybrid connectivity, observability, and platform reliability; technical leadership and delivery experience.
Google Cloud, Salesforce, Terraform, Kubernetes, CI/CD, VPC Service Controls, Cloud Armor, Cloud IDS, Cloud NGFW Enterprise, Network Connectivity Center, Partner Interconnect, HA VPN, IAM, Organization Policies
2mo
Save
Mark Applied
Hide
Site Reliability Engineer (Google Cloud Platf
United States
RemoteFull Time
Undefined
Undefined: London-based digital product studio building websites and digital products for ambitious companies.
5+ YOEU.S. citizen with active Secret clearance; Bachelor's in CS or related; 5+ years cloud/SRE experience with 3+ years on GCP; experience with GCP security, NIST/FedRAMP/CMMC, Terraform, EM/CI tools, and Python/Go/Bash.
Google Chronicle, Security Command Center, Cloud KMS, Cloud HSM, Cloud EKM, Cloud Audit Logs, BigQuery, Cloud Monitoring, Cloud Logging, Cloud Build, Cloud Deploy, Terraform, Infrastructure Manager, Cloud Foundation Toolkit, Binary Authorization, Artifact Registry, YARA-L, Cloud Armor, VPC Service Controls, Access Context Manager, Assured Workloads, Workload Identity Federation, Shared VPC, Cloud NGFW, Private Google Access, Private Service Connect, Access Approval, Access Transparency, BeyondCorp Enterprise, IAP, Python, Go, Bash
2w
Save
Mark Applied
Hide
Senior Site Reliability Engineer Platform Private Cloud Engineer
San Jose, California, United States
$94k-$130k/yr OnsiteFull Time
Tata Consultancy Services
Tata Consultancy ServicesBSE: 532540: Global leader in IT services, consulting, and business solutions.
7+ YOERequires 7+ years designing and operating enterprise or cloud environments, private cloud and Kubernetes expertise, scripting, IaC tools, Unix/Linux knowledge, and a CS or engineering degree.
VMware, AWS, GCP, Kubernetes, Helm, ArgoCD, Python, Bash, Ruby, Scala, Ansible, Terraform, Unix, Linux
1mo
Save
Mark Applied
Hide
Site Reliability Engineer
Charlotte, North Carolina, United States
HybridFull Time
Electrolux Group
Electrolux GroupNasdaq Stockholm: ELUX B: Global home appliance manufacturer reinventing taste, care, and wellbeing.
6+ YOE6+ years in infrastructure/site reliability/cloud engineering; experience with cloud platforms, IaC, CI/CD, observability, troubleshooting, and strong collaboration skills.
Microsoft Azure, AWS, Google Cloud Platform, Akamai CDN, Terraform, CloudFormation, Ansible, Puppet, Chef, Microsoft Azure DevOps, GitHub, Argo CD
2mo
Save
Mark Applied
Hide
Senior Reliability Engineer
Washington, District of Columbia, United States
OnsiteFull Time
Barbaricum
Barbaricum: Government contracting firm providing intelligence, analytics, and engineering support.
10+ YOE10+ years SRE/systems administration experience, Bachelor\u0002s in CS/IT/related (Master's preferred), DoD Secret clearance, expertise in monitoring, automation, cloud (AWS, Microsoft Azure, Google Cloud), scripting (Python, Shell, PowerShell), and configuration management tools.
Ansible, Puppet, Chef, Python, Shell, Microsoft PowerShell, AWS, Microsoft Azure, Google Cloud, Windows, Linux
1mo
Save
Mark Applied
Hide
Sr. Cloud Operations Reliability Engineer (SRE)
Georgia, United States
RemoteFull Time
NextGen Healthcare
NextGen Healthcare: Private U.S. healthcare technology providing cloud EHR, practice management, and patient-care software to ambulatory practices.
10+ YOE10+ years in cloud operations/SRE with GCP/AWS, SLO/SLI, incident response, IaC, Kubernetes, observability, and mentoring experience.
Google Cloud Platform (GCP), AWS, Terraform, Deployment Manager, CloudFormation, Kubernetes, Grafana, Prometheus, Cloud Monitoring, Datadog, New Relic, Python, Bash, Go, CI/CD
1mo
Save
Mark Applied
Hide
Alibaba Cloud-Cloud Infrastructure – Site Reliability Engineer (SRE)-Sunnyvale
Sunnyvale, California, United States
$104k-$171k/yr OnsiteFull Time
Alibaba Cloud
Alibaba CloudNYSE, HKEX: BABA, 9988: Global cloud computing and data intelligence service provider.
2+ YOE2+ years in distributed systems reliability engineering; high-availability architecture, Kafka/RocketMQ, Kubernetes, automation, and proficiency in Python, Go, or Java required. Bachelor's degree listed.
RocketMQ, Kafka, Kubernetes, K8s, Java, Go, Python, Shell, Terraform, Helm, Operator
2mo
Save
Mark Applied
Hide
Platform / Reliability Engineer
United States or India
RemoteFull Time
Finn
Finn: Finn is a private Indian enterprise voice-AI platform helping businesses automate inbound and outbound customer calls.
4+ YOERequires 4+ years in platform, SRE, or infrastructure engineering and experience with cloud platforms, containers, infrastructure as code, observability, deployment, reliability, and incident response.
GCP, AWS, CI/CD, containers, IaC, SOC 2, HIPAA
3w
Save
Mark Applied
Hide
Systems Reliability Engineer
Overland Park or Atlanta or Frisco or Bellevue
$84k-$151k/yr OnsiteFull Time
T-Mobile
T-MobileNASDAQ: TMUS: The Un-carrier providing wireless and home internet services.
2+ YOEBachelor's degree required; 2–4+ years preferred. Requires DevOps, cloud, automation, monitoring, scripting, APIs, cybersecurity, and reliability engineering experience, plus U.S. work authorization.
C, C#, Java, Perl, Python, Go, Jenkins, CloudBees, Ansible, Chef, Puppet, Docker, Kubernetes, AppDynamics, Splunk, Microsoft Graph API, REST API, Microsoft Power Apps, Microsoft Power Automate, Microsoft Entra, SailPoint, ServiceNow, Azure, Microsoft Azure DevOps Pipelines, Linux, Shell
4w
Save
Mark Applied
Hide
Senior Site Reliability Engineer - Cloud
Santa Clara, California, United States
$168k-$265k/yr OnsiteFull Time
NVIDIA
NVIDIANASDAQ: NVDA: Computing platform for AI and accelerated graphics.
8+ YOEMS/BS or equivalent experience, 8+ years supporting live-site production, SRE on-call experience, strong Kubernetes and Python skills, Akamai/CDN and AWS experience, incident management and automation focus.
Akamai Edge Redirector Cloudlets, Akamai Forward Rewrite Cloudlets, Akamai Cloudlets Policy Manager, Akamai CDN, WAF, AWS, Kubernetes, Python
4w
Save
Mark Applied
Hide
Senior Site Reliability Engineer - Cloud
Santa Clara, California, United States
$168k-$265k/yr OnsiteFull Time
NVIDIA
NVIDIANASDAQ: NVDA: Computing platform for AI and accelerated graphics.
8+ YOE8+ years supporting live-site production environments; BS/MS or equivalent; strong Kubernetes, AWS, Python; Akamai/CDN and SRE on-call experience required.
Akamai Edge Redirector Cloudlets, Akamai Forward Rewrite Cloudlets, Akamai Cloudlets Policy Manager, Akamai CDN, WAF, AWS, Kubernetes, Python
2w
Save
Mark Applied
Hide
Senior Site Reliability Engineer
Idaho, United States
$117k-$209k/yr RemoteFull Time
Autodesk
AutodeskNASDAQ: ADSK: Global provider of software for design, engineering, and manufacturing.
7+ YOEBachelor's degree or equivalent practical experience and 7+ years in SRE, software, platform, cloud infrastructure, or production operations; experience with cloud platforms, automation, IaC, CI/CD, and reliability engineering.
AWS, Azure, Python, Go, Java, PowerShell, Bash, Infrastructure as Code, CI/CD, Splunk, Dynatrace, Datadog, CloudWatch, Kubernetes
2mo
Save
Mark Applied
Hide
Senior Site Reliability Engineer
United States
$81k-$187k/yr OnsiteFull Time
Oracle Health
Oracle Health: Healthcare information-technology division providing cloud-based clinical, financial, interoperability, and population-health solutions to providers, payers, and public-health organizations.
3+ YOE3+ years SRE or reliability engineering experience; design and operate cloud-native infrastructure on Oracle Cloud; incident response, automation, observability; strong Python skills; security clearance may be required.
OCI - Kubernetes Engine, Oracle Cloud Infrastructure (OCI), Python