309 cloud reliability engineer jobs at 165 companies in California

2mo
Save
Mark Applied
Hide
Staff Cloud Reliability Engineer
Irvine or Los Angeles
$180k-$200k/yr OnsiteFull Time
Viant Technology
Viant TechnologyNasdaq Global Select Market: DSP: Public AI-powered advertising platform helping advertisers buy and measure connected-TV and open-internet campaigns programmatically.
8+ YOE8+ years in DevOps/SRE, 3+ years Linux, cloud (AWS/Google), serverless (AWS Lambda/Google Cloud Functions), Docker/Kubernetes, Terraform, CI/CD (GitHub Actions), Python or Go, SQL/BigQuery; participate in on-call rotation.
Linux, AWS, Google, AWS Lambda, Google Cloud Functions, Docker, Kubernetes, Terraform, GitHub Actions, Python, GoLang, SQL, Google BigQuery
3w
Save
Mark Applied
Hide
Principal Engineer, Cloud Site Reliability Engineering
Santa Clara, California, United States
$272k-$431k/yr OnsiteFull Time
NVIDIA
NVIDIANASDAQ: NVDA: Computing platform for AI and accelerated graphics.
15+ YOEBS/MS in engineering or computer science, 15+ years of systems software development including 1+ year in AI, cloud infrastructure experience, and strong Java, Python, Shell, distributed systems, and database skills.
Java, Python, Shell, REST APIs, MySQL, Cassandra, MongoDB, Elasticsearch, Docker, Virtual Machines, OpenStack, Kubernetes, Chef, Puppet, Hadoop, Ceph, SwiftStack, LXC, Git, Perforce, JFrog, Kafka, Windows, Linux, Android
1mo
Save
Mark Applied
Hide
Cloud Site Reliability Engineer - DCS Cloud
San Jose, California, United States
OnsiteFull Time
ByteDance
ByteDance: Global technology specializing in AI-powered content platforms.
2+ YOEBachelor's degree in CS or related,2+ years in Linux operations/SRE/DevOps,programming in Go/Python/C++,cloud and reliability practices experience,strong troubleshooting and communication skills.
Go, Python, C++, Linux, OCI, AWS, Azure, GCP, KVM, QEMU, Docker, Kubernetes, containerd, cgroups, namespaces, CUDA, MIG
1mo
Save
Mark Applied
Hide
Staff Network Reliability Engineer (Cloud Operations)
Mountain View, California, United States
HybridFull Time
Skylo Technologies
Skylo Technologies: Private telecommunications providing satellite connectivity for smartphones, vehicles, and IoT devices where cellular networks are unavailable.
8+ YOE8+ years cloud/infrastructure/SRE experience with Kubernetes, hybrid cloud operations, observability, database and storage reliability, GitOps, and on-call ownership in 24x7 environments.
Kubernetes, GKE, GCP, kubectl, Pub/Sub, Cloud SQL, Prometheus, VictoriaMetrics, Grafana, OpenTelemetry, PostgreSQL, Redis, ArgoCD, Helm, Terraform, Ansible, Ceph, Rook, Harvester, KubeVirt, KVM, Loki, ELK, Flux CD, Go, Python, BGP, VXLAN, EVPN
2w
Save
Mark Applied
Hide
Senior Site Reliability Engineer Platform Private Cloud Engineer
San Jose, California, United States
$94k-$130k/yr OnsiteFull Time
Tata Consultancy Services
Tata Consultancy ServicesBSE: 532540: Global leader in IT services, consulting, and business solutions.
7+ YOERequires 7+ years designing and operating enterprise or cloud environments, private cloud and Kubernetes expertise, scripting, IaC tools, Unix/Linux knowledge, and a CS or engineering degree.
VMware, AWS, GCP, Kubernetes, Helm, ArgoCD, Python, Bash, Ruby, Scala, Ansible, Terraform, Unix, Linux
1mo
Save
Mark Applied
Hide
Alibaba Cloud-Cloud Infrastructure – Site Reliability Engineer (SRE)-Sunnyvale
Sunnyvale, California, United States
$104k-$171k/yr OnsiteFull Time
Alibaba Cloud
Alibaba CloudNYSE, HKEX: BABA, 9988: Global cloud computing and data intelligence service provider.
2+ YOE2+ years in distributed systems reliability engineering; high-availability architecture, Kafka/RocketMQ, Kubernetes, automation, and proficiency in Python, Go, or Java required. Bachelor's degree listed.
RocketMQ, Kafka, Kubernetes, K8s, Java, Go, Python, Shell, Terraform, Helm, Operator
4w
Save
Mark Applied
Hide
Senior Site Reliability Engineer - Cloud
Santa Clara, California, United States
$168k-$265k/yr OnsiteFull Time
NVIDIA
NVIDIANASDAQ: NVDA: Computing platform for AI and accelerated graphics.
8+ YOE8+ years supporting live-site production environments; BS/MS or equivalent; strong Kubernetes, AWS, Python; Akamai/CDN and SRE on-call experience required.
Akamai Edge Redirector Cloudlets, Akamai Forward Rewrite Cloudlets, Akamai Cloudlets Policy Manager, Akamai CDN, WAF, AWS, Kubernetes, Python
4w
Save
Mark Applied
Hide
Senior Site Reliability Engineer - Cloud
Santa Clara, California, United States
$168k-$265k/yr OnsiteFull Time
NVIDIA
NVIDIANASDAQ: NVDA: Computing platform for AI and accelerated graphics.
8+ YOEMS/BS or equivalent experience, 8+ years supporting live-site production, SRE on-call experience, strong Kubernetes and Python skills, Akamai/CDN and AWS experience, incident management and automation focus.
Akamai Edge Redirector Cloudlets, Akamai Forward Rewrite Cloudlets, Akamai Cloudlets Policy Manager, Akamai CDN, WAF, AWS, Kubernetes, Python
1mo
Save
Mark Applied
Hide
Cloud Engineer
Palo Alto, California, United States
$125k-$200k/yr OnsiteFull Time
Sage Care
Sage Care: Private healthcare AI platform helping health systems automate patient support, triage, scheduling, and provider matching.
4+ YOE4+ years DevOps/SRE experience with GCP, Kubernetes (GKE), Terraform, CI/CD, and Bazel; strong networking, IAM, cloud security, and production reliability skills.
GCP, Terraform, Kubernetes, GKE, Bazel, CI/CD
1mo
Save
Mark Applied
Hide
Site Reliability Engineer
California, United States
OnsiteFull Time
Arena Intelligence
Arena Intelligence: AI model evaluation platform serving enterprises, AI labs, and independent researchers in real-world workflows.
6+ YOE6+ years backend engineering with distributed systems, proficiency in Go or Rust, experience with LLM provider APIs, cloud (AWS/GCP), Kubernetes, Terraform, Postgres, and Redis.
Go, Rust, OpenAI, Anthropic, Google, AWS, GCP, Kubernetes, Terraform, Postgres, Redis, Bifrost, Kong, Envoy, Tyk, Stripe, Metronome, Orb, vLLM, LiteLLM, LangChain
1mo
Save
Mark Applied
Hide
Lead Site Reliability Engineer
Los Angeles, California, United States
$140k-$199k/yr HybridFull Time
Green Dot Corporation
Green Dot CorporationNYSE: GDOT: Public U.S. fintech bank holding providing banking and payment services to consumers and businesses.
7+ YOE7+ years in release/reliability engineering, cloud platform experience (AWS/Azure/GCP), automated deployment and observability proficiency, scripting with PowerShell/Bash/Python, excellent troubleshooting and communication skills.
AWS, Azure, GCP, PowerShell, Bash, Python
3w
Save
Mark Applied
Hide
Site Reliability Engineer (SRE)
Santa Clara, California, United States
$50-$60/hr RemoteContract
ServiceNow
ServiceNowNYSE: NOW: Enterprise software providing cloud-based workflow automation platforms.
3+ YOEBachelor's degree in computer science or related field; 3+ years in site reliability engineering; 2+ years with AWS and cloud automation; Kubernetes, Linux, Terraform, networking, GitOps, monitoring, and customer support experience.
AWS, Kubernetes, Helm, Linux, Terraform, GitOps, Prometheus, Grafana, Bazel, CueLang, Version Control, Okta, Snowflake, Google
2mo
Save
Mark Applied
Hide
Senior Site Reliability Engineer
Oakland, California, United States
$175k-$210k/yr HybridFull Time
Fivetran
Fivetran: Automated data movement and integration platform for organizations.
5+ YOE5+ years SaaS experience; managed Kubernetes, cloud platforms (AWS/GCP/Azure), Terraform/Ansible/ArgoCD; Python/Shell scripting, Linux admin, PostgreSQL; incident response and reliability engineering experience.
Kubernetes, EKS, AKS, GKE, PostgreSQL, ArgoCD, Terraform, Ansible, Python, Shell, Go, Java, AWS, GCP, Azure, Grafana, Buildkite, Temporal, Pulumi, Linux, VPN, PrivateLink, Private Service Connect (GCP)
1mo
Save
Mark Applied
Hide
Site Reliability Engineer
San Francisco, California, United States
HybridFull Time
Runloop AI
Runloop AI: Runloop AI provides AI infrastructure, secure code sandboxes, and evaluation tools for developers building software-engineering agents.
5+ YOE5+ years software engineering experience with 3+ years in SRE/DevOps, strong Python or Go skills, containerization, cloud infra, monitoring, networking, Linux administration, on‑call and incident management.
AWS, GCP, Azure, Grafana, Prometheus, Datadog, Python, Go, Docker, Kubernetes, Terraform, Pulumi, Sentry, RUM, CI/CD
3w
Save
Mark Applied
Hide
Site Reliability Engineer (SRE)
San Francisco, California, United States
$350k-$475k/yr OnsiteFull Time
Thinking Machines Lab
Thinking Machines Lab: Private AI research and product building customizable multimodal systems for researchers and the wider public.
Experience in distributed systems/cloud/site reliability, software automation for reliability, incident response and postmortems, strong communication and coordination skills.
Tinker, Kubernetes, LoRA, CI/CD
2mo
Save
Mark Applied
Hide
Principal Site Reliability Engineer, Google Cloud
Atlanta or Milpitas
$240k-$250k/yr HybridFull Time
Saviynt
Saviynt: Private enterprise software providing AI-powered identity security and access governance for global enterprises and government institutions.
9+ YOE9+ years in platform/infra/SRE roles, deep Kubernetes and GCP expertise, strong Go and Python skills, experience with CI/CD, event-driven systems, observability, distributed systems, and building shared platform services.
Go (Golang), Python, Kubernetes, GCP, AWS, Azure, Kafka, RMQ, NATS, Google Pub/Sub, GitLab CI, ArgoCD, Prometheus, Grafana, ELK stack, Datadog, Envoy, Istio, MySQL, PostgresSQL
3w
Save
Mark Applied
Hide
Site Reliability Engineer
United States or California
$110k-$253k/yr RemoteFull Time
Veeam Government Solutions
Veeam Government Solutions: The Data and AI Trust.
7+ YOE7+ years software engineering with 3+ years SRE/platform experience, government/regulated cloud familiarity, Azure and IaC experience, programming (TypeScript/JS, Go, Java, C#), observability and CI/CD experience.
Microsoft TFS, Azure DevOps, Git, BitBucket, Azure, Azure Government, Azure Monitor, Application Insights, Cosmos Db, Azure Functions, ARM templates, AWS CloudFormation, Terraform, Terragrunt, Pulumi, Serverless Framework, Kubernetes, Prometheus, Grafana, OpenTelemetry, Elastic Stack, GitHub Actions, GitLab CI, ArgoCD, FluxCD, Dagger, TypeScript, JavaScript, Go, Java, C#
1mo
Save
Mark Applied
Hide
Site Reliability Engineer
Santa Clara, California, United States
$230k-$250k/yr OnsiteFull Time
Forward Networks
Forward Networks: Private network-software helping enterprises and government agencies model, secure, and automate complex hybrid networks.
6+ YOE6+ years SRE/DevOps experience in SaaS/cloud, strong networking fundamentals, Kubernetes, observability (Prometheus/Grafana/Datadog/Splunk), Python/Bash automation, cloud and IaC (AWS/GCP/Azure, Terraform/Ansible), and incident response ownership.
Kubernetes, Prometheus, Grafana, Datadog, Splunk, Python, Bash, AWS, GCP, Azure, Terraform, Ansible
3w
Save
Mark Applied
Hide
Principal Site Reliability Engineer (Hybrid)
Merrimack or San Diego
$118k-$201k/yr HybridFull Time
BAE Systems
BAE SystemsLSE: BA.: Global defense, aerospace, and security technology.
4+ YOERequires 4–6+ years of site reliability engineering, Juniper networking, cloud technologies, automation, storage, virtualization, and security clearance eligibility; Security+ required or obtainable within 90 days.
Juniper, Ansible, Helm Charts, NFS, JDFS, Ceph, S3, VMware, Open Stack, Azure Stack, Kubernetes, Terraform
2mo
Save
Mark Applied
Hide
Senior Software Engineer- Site Reliability Engineering (SRE)
California or Virginia or Maryland or United States or Reston or San Diego
$149k-$202k/yr RemoteFull Time
Noctua Technology
Noctua Technology: Building a safe and equitable future through technology.
5+ YOE5+ years SRE/cloud engineering experience; strong software engineering, IaC (Terraform/CloudFormation), Docker/Kubernetes, Python/Bash/Go, CI/CD, cloud security, and ability to obtain Secret clearance.
Terraform, CloudFormation, Docker, Kubernetes, Python, Bash, Go, Google Cloud, AWS

Explore Jobs

Expand Your Job Search