538 cloud site reliability engineer jobs at 293 companies in United States

1mo
Save
Mark Applied
Hide
Site Reliability Engineer
Charlotte, North Carolina, United States
HybridFull Time
Electrolux Group
Electrolux GroupNasdaq Stockholm: ELUX B: Global home appliance manufacturer reinventing taste, care, and wellbeing.
6+ YOE6+ years in infrastructure/site reliability/cloud engineering; experience with cloud platforms, IaC, CI/CD, observability, troubleshooting, and strong collaboration skills.
Microsoft Azure, AWS, Google Cloud Platform, Akamai CDN, Terraform, CloudFormation, Ansible, Puppet, Chef, Microsoft Azure DevOps, GitHub, Argo CD
3w
Save
Mark Applied
Hide
Principal Engineer, Cloud Site Reliability Engineering
Santa Clara, California, United States
$272k-$431k/yr OnsiteFull Time
NVIDIA
NVIDIANASDAQ: NVDA: Computing platform for AI and accelerated graphics.
15+ YOEBS/MS in engineering or computer science, 15+ years of systems software development including 1+ year in AI, cloud infrastructure experience, and strong Java, Python, Shell, distributed systems, and database skills.
Java, Python, Shell, REST APIs, MySQL, Cassandra, MongoDB, Elasticsearch, Docker, Virtual Machines, OpenStack, Kubernetes, Chef, Puppet, Hadoop, Ceph, SwiftStack, LXC, Git, Perforce, JFrog, Kafka, Windows, Linux, Android
1mo
Save
Mark Applied
Hide
Cloud Site Reliability Engineer - DCS Cloud
San Jose, California, United States
OnsiteFull Time
ByteDance
ByteDance: Global technology specializing in AI-powered content platforms.
2+ YOEBachelor's degree in CS or related,2+ years in Linux operations/SRE/DevOps,programming in Go/Python/C++,cloud and reliability practices experience,strong troubleshooting and communication skills.
Go, Python, C++, Linux, OCI, AWS, Azure, GCP, KVM, QEMU, Docker, Kubernetes, containerd, cgroups, namespaces, CUDA, MIG
3w
Save
Mark Applied
Hide
Senior Site Reliability Engineer - Cloud
Santa Clara, California, United States
$168k-$265k/yr OnsiteFull Time
NVIDIA
NVIDIANASDAQ: NVDA: Computing platform for AI and accelerated graphics.
8+ YOE8+ years supporting live-site production environments; BS/MS or equivalent; strong Kubernetes, AWS, Python; Akamai/CDN and SRE on-call experience required.
Akamai Edge Redirector Cloudlets, Akamai Forward Rewrite Cloudlets, Akamai Cloudlets Policy Manager, Akamai CDN, WAF, AWS, Kubernetes, Python
6d
Save
Mark Applied
Hide
Senior Site Reliability Engineer
Nashville, Tennessee, United States
$81k-$187k/yr OnsiteFull Time
Oracle Corporation
Oracle CorporationNYSE: ORCL: Cloud infrastructure and enterprise software solutions provider.
3+ YOEBachelor’s degree in Computer Science or equivalent experience; 3+ years in Site Reliability Engineering, DevOps, or Systems Engineering; cloud operations, incident management, automation, programming, and infrastructure tooling experience.
AWS, Azure, GCP, OCI, Chef, Ansible, Jenkins, Terraform, Docker, RESTful APIs, CI/CD, Agile
3w
Save
Mark Applied
Hide
Senior Site Reliability Engineer - Cloud
Santa Clara, California, United States
$168k-$265k/yr OnsiteFull Time
NVIDIA
NVIDIANASDAQ: NVDA: Computing platform for AI and accelerated graphics.
8+ YOEMS/BS or equivalent experience, 8+ years supporting live-site production, SRE on-call experience, strong Kubernetes and Python skills, Akamai/CDN and AWS experience, incident management and automation focus.
Akamai Edge Redirector Cloudlets, Akamai Forward Rewrite Cloudlets, Akamai Cloudlets Policy Manager, Akamai CDN, WAF, AWS, Kubernetes, Python
3w
Save
Mark Applied
Hide
Site Reliability Engineer (SRE)
San Francisco, California, United States
$350k-$475k/yr OnsiteFull Time
Thinking Machines Lab
Thinking Machines Lab: Private AI research and product building customizable multimodal systems for researchers and the wider public.
Experience in distributed systems/cloud/site reliability, software automation for reliability, incident response and postmortems, strong communication and coordination skills.
Tinker, Kubernetes, LoRA, CI/CD
2mo
Save
Mark Applied
Hide
Site Reliability Engineer (Google Cloud Platf
United States
RemoteFull Time
Undefined
Undefined: London-based digital product studio building websites and digital products for ambitious companies.
5+ YOEU.S. citizen with active Secret clearance; Bachelor's in CS or related; 5+ years cloud/SRE experience with 3+ years on GCP; experience with GCP security, NIST/FedRAMP/CMMC, Terraform, EM/CI tools, and Python/Go/Bash.
Google Chronicle, Security Command Center, Cloud KMS, Cloud HSM, Cloud EKM, Cloud Audit Logs, BigQuery, Cloud Monitoring, Cloud Logging, Cloud Build, Cloud Deploy, Terraform, Infrastructure Manager, Cloud Foundation Toolkit, Binary Authorization, Artifact Registry, YARA-L, Cloud Armor, VPC Service Controls, Access Context Manager, Assured Workloads, Workload Identity Federation, Shared VPC, Cloud NGFW, Private Google Access, Private Service Connect, Access Approval, Access Transparency, BeyondCorp Enterprise, IAP, Python, Go, Bash
2w
Save
Mark Applied
Hide
Site Reliability Engineer (SRE)
Santa Clara, California, United States
$50-$60/hr RemoteContract
ServiceNow
ServiceNowNYSE: NOW: Enterprise software providing cloud-based workflow automation platforms.
3+ YOEBachelor's degree in computer science or related field; 3+ years in site reliability engineering; 2+ years with AWS and cloud automation; Kubernetes, Linux, Terraform, networking, GitOps, monitoring, and customer support experience.
AWS, Kubernetes, Helm, Linux, Terraform, GitOps, Prometheus, Grafana, Bazel, CueLang, Version Control, Okta, Snowflake, Google
3w
Save
Mark Applied
Hide
Site Reliability Engineer
Westminster or Lake Oswego
$106k-$145k/yr OnsiteFull Time
Trimble
TrimbleNASDAQ: TRMB: Providing technology solutions that connect the physical and digital worlds.
3+ YOE3+ years SRE/DevOps/Cloud Engineering experience, strong Microsoft Azure and Terraform skills, CI/CD and scripting (PowerShell, Bash), GitHub Actions/Azure DevOps, New Relic experience, enterprise cloud proficiency.
Microsoft Azure, Terraform, PowerShell, Bash, GitHub Actions, Azure DevOps, New Relic
2mo
Save
Mark Applied
Hide
Site Reliability Engineer
Washington, District of Columbia, United States
$112k-$179k/yr OnsiteFull Time
Peraton
Peraton: Provider of mission-critical national security technologies and services.
7+ YOERequires active TS/SCI clearance, bachelor's in CS/IT or equivalent, 7+ years software engineering/DevOps experience, cloud certifications, CISSP/CASP+ preferred, DoD 8140/8570 compliance, strong Linux and cloud-native automation experience.
Kubernetes, Rancher, Helm, Docker, cilium, Rook, Ceph, MinIO, S3, PortWorz, Ansible, Terraform, Python, PowerShell, Linux, Desired State Configuration
2mo
Save
Mark Applied
Hide
Site Reliability Engineer
United States
RemoteFull Time
Supabase
Supabase: Developer platform providing Postgres databases, authentication, storage, realtime, REST APIs, and edge functions for application developers.
7+ YOE7+ years in SRE/production engineering, experience shaping SRE practices, defining and operationalizing SLOs/SLIs, incident response and postmortems, software engineering mindset, cloud infra (AWS) and IaC (Pulumi/Terraform/CDK).
Postgres, AWS, Pulumi, Terraform, CDK, Kubernetes, OpenTelemetry, VictoriaMetrics, Grafana, DORA metrics
2mo
Save
Mark Applied
Hide
Site Reliability Engineer
Vienna, Virginia, United States
RemoteFull Time
Knexus
Knexus: Private AI research and engineering delivering secure, tested systems and data-science solutions to U.S. government agencies.
6+ YOE6+ years in infrastructure engineering; strong cloud expertise (GCP/AWS/Azure); security controls (NIST 800-53/800-171); DoD Cloud SRG; ATO/SSP experience; GCP certification desired; US citizen eligible for security clearance.
Google Cloud Platform, Amazon Web Services, Microsoft Azure, Kubernetes, IAM, SSP, ATO, NIST 800-53/800-171, DoD Cloud SRG, Google Cloud Professional certifications
2w
Save
Mark Applied
Hide
Senior Site Reliability Engineer Platform Private Cloud Engineer
San Jose, California, United States
$94k-$130k/yr OnsiteFull Time
Tata Consultancy Services
Tata Consultancy ServicesBSE: 532540: Global leader in IT services, consulting, and business solutions.
7+ YOERequires 7+ years designing and operating enterprise or cloud environments, private cloud and Kubernetes expertise, scripting, IaC tools, Unix/Linux knowledge, and a CS or engineering degree.
VMware, AWS, GCP, Kubernetes, Helm, ArgoCD, Python, Bash, Ruby, Scala, Ansible, Terraform, Unix, Linux
2w
Save
Mark Applied
Hide
Senior Site Reliability Engineer
Idaho, United States
$117k-$209k/yr RemoteFull Time
Autodesk
AutodeskNASDAQ: ADSK: Global provider of software for design, engineering, and manufacturing.
7+ YOEBachelor's degree or equivalent practical experience and 7+ years in SRE, software, platform, cloud infrastructure, or production operations; experience with cloud platforms, automation, IaC, CI/CD, and reliability engineering.
AWS, Azure, Python, Go, Java, PowerShell, Bash, Infrastructure as Code, CI/CD, Splunk, Dynatrace, Datadog, CloudWatch, Kubernetes
2mo
Save
Mark Applied
Hide
Site Reliability Engineer
Austin, Texas, United States
$140k-$200k/yr OnsiteFull Time
Future Secure AI
Future Secure AI: Private enterprise AI building and operating bespoke AI-Workers that automate complex, high-stakes workflows.
5+ YOEHands-on Kubernetes, Terraform, and Helm experience; programming in Python/Go/Java/Bash/PowerShell/Ruby; SRE experience with on-call, incident response, SLIs/SLOs; cloud and CI/CD experience; 5+ years preferred.
Kubernetes, EKS, AKS, GKE, Terraform, Helm, SLIs, SLOs, SLAs, Python, Go, Java, Bash, PowerShell, Ruby, ArgoCD, CI/CD, GitOps, AWS, Azure, Google Cloud
1mo
Save
Mark Applied
Hide
Site Reliability Engineer
California, United States
OnsiteFull Time
Arena Intelligence
Arena Intelligence: AI model evaluation platform serving enterprises, AI labs, and independent researchers in real-world workflows.
6+ YOE6+ years backend engineering with distributed systems, proficiency in Go or Rust, experience with LLM provider APIs, cloud (AWS/GCP), Kubernetes, Terraform, Postgres, and Redis.
Go, Rust, OpenAI, Anthropic, Google, AWS, GCP, Kubernetes, Terraform, Postgres, Redis, Bifrost, Kong, Envoy, Tyk, Stripe, Metronome, Orb, vLLM, LiteLLM, LangChain
2mo
Save
Mark Applied
Hide
Site Reliability Engineer
Centennial, Colorado, United States
$110k-$145k/yr HybridFull Time
NBCUniversal
NBCUniversalNASDAQ: CMCSA: A leading global media and entertainment.
3+ YOEBachelor's in CS/IT or equivalent experience, 3+ years support experience, Linux/Unix and networking skills, scripting (Python, Bash, JavaScript, JSON, XML, YML), AWS and infrastructure automation (Cloud Formation, Terraform, Ansible, Chef), Git experience, and on-call participation.
Linux/Unix, Python, Bash, JavaScript, JSON, XML, YML, AWS, Cloud Formation, Terraform, Ansible, Chef, Git, Git CLI, Kubernetes, Docker, EKS, Microsoft Graph API, HEVC, AVC, AC3, AAC, MPEG Transport Streams, ATSC, HLS, CMAF, Zixi, SRT, RIST, 2022-7, 2110, SCTE35, SCTE104, SCTE224, SSAI, StatMUX
2mo
Save
Mark Applied
Hide
Site Reliability Engineer
London or Leeds or Riga or Lisbon or Australia or France or Ireland or Latvia or Portugal or United States
£78k-£118k/yr OnsiteFull Time
GoCardless
GoCardless: UK private fintech helping businesses collect and send recurring and one-off payments directly from customers’ bank accounts.
Experience with cloud infrastructure, IaC, container orchestration, CI/CD, monitoring, and programming (e.g., Python, Ruby, Golang). Comfortable with on-call rotations, incident management, automation, and collaborating with engineering teams.
Python, Ruby, Golang, Terraform, Terragrunt, Atlantis, AWS, GCP, Kubernetes, GKE, GitHub, GitHub Actions, ArgoCD, Grafana, Prometheus, Datadog
1mo
Save
Mark Applied
Hide
Site Reliability Engineer
San Francisco, California, United States
HybridFull Time
Runloop AI
Runloop AI: Runloop AI provides AI infrastructure, secure code sandboxes, and evaluation tools for developers building software-engineering agents.
5+ YOE5+ years software engineering experience with 3+ years in SRE/DevOps, strong Python or Go skills, containerization, cloud infra, monitoring, networking, Linux administration, on‑call and incident management.
AWS, GCP, Azure, Grafana, Prometheus, Datadog, Python, Go, Docker, Kubernetes, Terraform, Pulumi, Sentry, RUM, CI/CD