538 cloud site reliability engineer jobs at 293 companies in United States
1mo
Save
Mark Applied
Hide
1mo
Site Reliability Engineer
Charlotte, North Carolina, United States
HybridFull Time
Electrolux GroupNasdaq Stockholm: ELUX B: Global home appliance manufacturer reinventing taste, care, and wellbeing.
6+ YOE6+ years in infrastructure/site reliability/cloud engineering; experience with cloud platforms, IaC, CI/CD, observability, troubleshooting, and strong collaboration skills.
Microsoft Azure, AWS, Google Cloud Platform, Akamai CDN, Terraform, CloudFormation, Ansible, Puppet, Chef, Microsoft Azure DevOps, GitHub, Argo CD
Principal Engineer, Cloud Site Reliability Engineering
Santa Clara, California, United States
$272k-$431k/yrOnsiteFull Time
NVIDIANASDAQ: NVDA: Computing platform for AI and accelerated graphics.
15+ YOEBS/MS in engineering or computer science, 15+ years of systems software development including 1+ year in AI, cloud infrastructure experience, and strong Java, Python, Shell, distributed systems, and database skills.
ByteDance: Global technology specializing in AI-powered content platforms.
2+ YOEBachelor's degree in CS or related,2+ years in Linux operations/SRE/DevOps,programming in Go/Python/C++,cloud and reliability practices experience,strong troubleshooting and communication skills.
NVIDIANASDAQ: NVDA: Computing platform for AI and accelerated graphics.
8+ YOE8+ years supporting live-site production environments; BS/MS or equivalent; strong Kubernetes, AWS, Python; Akamai/CDN and SRE on-call experience required.
Oracle CorporationNYSE: ORCL: Cloud infrastructure and enterprise software solutions provider.
3+ YOEBachelor’s degree in Computer Science or equivalent experience; 3+ years in Site Reliability Engineering, DevOps, or Systems Engineering; cloud operations, incident management, automation, programming, and infrastructure tooling experience.
Thinking Machines Lab: Private AI research and product building customizable multimodal systems for researchers and the wider public.
Experience in distributed systems/cloud/site reliability, software automation for reliability, incident response and postmortems, strong communication and coordination skills.
Undefined: London-based digital product studio building websites and digital products for ambitious companies.
5+ YOEU.S. citizen with active Secret clearance; Bachelor's in CS or related; 5+ years cloud/SRE experience with 3+ years on GCP; experience with GCP security, NIST/FedRAMP/CMMC, Terraform, EM/CI tools, and Python/Go/Bash.
3+ YOEBachelor's degree in computer science or related field; 3+ years in site reliability engineering; 2+ years with AWS and cloud automation; Kubernetes, Linux, Terraform, networking, GitOps, monitoring, and customer support experience.
AWS, Kubernetes, Helm, Linux, Terraform, GitOps, Prometheus, Grafana, Bazel, CueLang, Version Control, Okta, Snowflake, Google
Peraton: Provider of mission-critical national security technologies and services.
7+ YOERequires active TS/SCI clearance, bachelor's in CS/IT or equivalent, 7+ years software engineering/DevOps experience, cloud certifications, CISSP/CASP+ preferred, DoD 8140/8570 compliance, strong Linux and cloud-native automation experience.
Supabase: Developer platform providing Postgres databases, authentication, storage, realtime, REST APIs, and edge functions for application developers.
7+ YOE7+ years in SRE/production engineering, experience shaping SRE practices, defining and operationalizing SLOs/SLIs, incident response and postmortems, software engineering mindset, cloud infra (AWS) and IaC (Pulumi/Terraform/CDK).
Knexus: Private AI research and engineering delivering secure, tested systems and data-science solutions to U.S. government agencies.
6+ YOE6+ years in infrastructure engineering; strong cloud expertise (GCP/AWS/Azure); security controls (NIST 800-53/800-171); DoD Cloud SRG; ATO/SSP experience; GCP certification desired; US citizen eligible for security clearance.
Google Cloud Platform, Amazon Web Services, Microsoft Azure, Kubernetes, IAM, SSP, ATO, NIST 800-53/800-171, DoD Cloud SRG, Google Cloud Professional certifications
Senior Site Reliability Engineer Platform Private Cloud Engineer
San Jose, California, United States
$94k-$130k/yrOnsiteFull Time
Tata Consultancy ServicesBSE: 532540: Global leader in IT services, consulting, and business solutions.
7+ YOERequires 7+ years designing and operating enterprise or cloud environments, private cloud and Kubernetes expertise, scripting, IaC tools, Unix/Linux knowledge, and a CS or engineering degree.
AutodeskNASDAQ: ADSK: Global provider of software for design, engineering, and manufacturing.
7+ YOEBachelor's degree or equivalent practical experience and 7+ years in SRE, software, platform, cloud infrastructure, or production operations; experience with cloud platforms, automation, IaC, CI/CD, and reliability engineering.
Future Secure AI: Private enterprise AI building and operating bespoke AI-Workers that automate complex, high-stakes workflows.
5+ YOEHands-on Kubernetes, Terraform, and Helm experience; programming in Python/Go/Java/Bash/PowerShell/Ruby; SRE experience with on-call, incident response, SLIs/SLOs; cloud and CI/CD experience; 5+ years preferred.
Arena Intelligence: AI model evaluation platform serving enterprises, AI labs, and independent researchers in real-world workflows.
6+ YOE6+ years backend engineering with distributed systems, proficiency in Go or Rust, experience with LLM provider APIs, cloud (AWS/GCP), Kubernetes, Terraform, Postgres, and Redis.
Runloop AI: Runloop AI provides AI infrastructure, secure code sandboxes, and evaluation tools for developers building software-engineering agents.
5+ YOE5+ years software engineering experience with 3+ years in SRE/DevOps, strong Python or Go skills, containerization, cloud infra, monitoring, networking, Linux administration, on‑call and incident management.