309 cloud reliability engineer jobs at 165 companies in California
2mo
Save
Mark Applied
Hide
2mo
Staff Cloud Reliability Engineer
Irvine or Los Angeles
$180k-$200k/yrOnsiteFull Time
Viant TechnologyNasdaq Global Select Market: DSP: Public AI-powered advertising platform helping advertisers buy and measure connected-TV and open-internet campaigns programmatically.
8+ YOE8+ years in DevOps/SRE, 3+ years Linux, cloud (AWS/Google), serverless (AWS Lambda/Google Cloud Functions), Docker/Kubernetes, Terraform, CI/CD (GitHub Actions), Python or Go, SQL/BigQuery; participate in on-call rotation.
Linux, AWS, Google, AWS Lambda, Google Cloud Functions, Docker, Kubernetes, Terraform, GitHub Actions, Python, GoLang, SQL, Google BigQuery
Principal Engineer, Cloud Site Reliability Engineering
Santa Clara, California, United States
$272k-$431k/yrOnsiteFull Time
NVIDIANASDAQ: NVDA: Computing platform for AI and accelerated graphics.
15+ YOEBS/MS in engineering or computer science, 15+ years of systems software development including 1+ year in AI, cloud infrastructure experience, and strong Java, Python, Shell, distributed systems, and database skills.
ByteDance: Global technology specializing in AI-powered content platforms.
2+ YOEBachelor's degree in CS or related,2+ years in Linux operations/SRE/DevOps,programming in Go/Python/C++,cloud and reliability practices experience,strong troubleshooting and communication skills.
Skylo Technologies: Private telecommunications providing satellite connectivity for smartphones, vehicles, and IoT devices where cellular networks are unavailable.
8+ YOE8+ years cloud/infrastructure/SRE experience with Kubernetes, hybrid cloud operations, observability, database and storage reliability, GitOps, and on-call ownership in 24x7 environments.
Senior Site Reliability Engineer Platform Private Cloud Engineer
San Jose, California, United States
$94k-$130k/yrOnsiteFull Time
Tata Consultancy ServicesBSE: 532540: Global leader in IT services, consulting, and business solutions.
7+ YOERequires 7+ years designing and operating enterprise or cloud environments, private cloud and Kubernetes expertise, scripting, IaC tools, Unix/Linux knowledge, and a CS or engineering degree.
Alibaba Cloud-Cloud Infrastructure – Site Reliability Engineer (SRE)-Sunnyvale
Sunnyvale, California, United States
$104k-$171k/yrOnsiteFull Time
Alibaba CloudNYSE, HKEX: BABA, 9988: Global cloud computing and data intelligence service provider.
2+ YOE2+ years in distributed systems reliability engineering; high-availability architecture, Kafka/RocketMQ, Kubernetes, automation, and proficiency in Python, Go, or Java required. Bachelor's degree listed.
NVIDIANASDAQ: NVDA: Computing platform for AI and accelerated graphics.
8+ YOE8+ years supporting live-site production environments; BS/MS or equivalent; strong Kubernetes, AWS, Python; Akamai/CDN and SRE on-call experience required.
Sage Care: Private healthcare AI platform helping health systems automate patient support, triage, scheduling, and provider matching.
4+ YOE4+ years DevOps/SRE experience with GCP, Kubernetes (GKE), Terraform, CI/CD, and Bazel; strong networking, IAM, cloud security, and production reliability skills.
Arena Intelligence: AI model evaluation platform serving enterprises, AI labs, and independent researchers in real-world workflows.
6+ YOE6+ years backend engineering with distributed systems, proficiency in Go or Rust, experience with LLM provider APIs, cloud (AWS/GCP), Kubernetes, Terraform, Postgres, and Redis.
Green Dot CorporationNYSE: GDOT: Public U.S. fintech bank holding providing banking and payment services to consumers and businesses.
7+ YOE7+ years in release/reliability engineering, cloud platform experience (AWS/Azure/GCP), automated deployment and observability proficiency, scripting with PowerShell/Bash/Python, excellent troubleshooting and communication skills.
3+ YOEBachelor's degree in computer science or related field; 3+ years in site reliability engineering; 2+ years with AWS and cloud automation; Kubernetes, Linux, Terraform, networking, GitOps, monitoring, and customer support experience.
AWS, Kubernetes, Helm, Linux, Terraform, GitOps, Prometheus, Grafana, Bazel, CueLang, Version Control, Okta, Snowflake, Google
Runloop AI: Runloop AI provides AI infrastructure, secure code sandboxes, and evaluation tools for developers building software-engineering agents.
5+ YOE5+ years software engineering experience with 3+ years in SRE/DevOps, strong Python or Go skills, containerization, cloud infra, monitoring, networking, Linux administration, on‑call and incident management.
Thinking Machines Lab: Private AI research and product building customizable multimodal systems for researchers and the wider public.
Experience in distributed systems/cloud/site reliability, software automation for reliability, incident response and postmortems, strong communication and coordination skills.
Saviynt: Private enterprise software providing AI-powered identity security and access governance for global enterprises and government institutions.
9+ YOE9+ years in platform/infra/SRE roles, deep Kubernetes and GCP expertise, strong Go and Python skills, experience with CI/CD, event-driven systems, observability, distributed systems, and building shared platform services.
BAE SystemsLSE: BA.: Global defense, aerospace, and security technology.
4+ YOERequires 4–6+ years of site reliability engineering, Juniper networking, cloud technologies, automation, storage, virtualization, and security clearance eligibility; Security+ required or obtainable within 90 days.