112 cloud reliability engineer jobs at 60 companies in Scotts Valley, CA

4w
Save
Mark Applied
Hide
Cloud Site Reliability Engineer - DCS Cloud
San Jose, California, United States
OnsiteFull Time
ByteDance
ByteDance: Developing AI-driven content platforms and mobile applications.
2+ YOEBachelor's degree in CS or related,2+ years in Linux operations/SRE/DevOps,programming in Go/Python/C++,cloud and reliability practices experience,strong troubleshooting and communication skills.
Go, Python, C++, Linux, OCI, AWS, Azure, GCP, KVM, QEMU, Docker, Kubernetes, containerd, cgroups, namespaces, CUDA, MIG
1mo
Save
Mark Applied
Hide
Senior Reliability Engineer, DGX Cloud
Santa Clara or United States
$168k-$334k/yr HybridFull Time
NVIDIA
NVIDIANASDAQ: NVDA: Designs GPU-accelerated computing and artificial intelligence hardware.
10+ YOE10+ years running large-scale production systems, strong software engineering (Go/Python), SLO program experience, incident response leadership, chaos engineering and failure-injection expertise, ability to influence across teams.
Go, Python, Prometheus, OpenTelemetry, Grafana, PagerDuty, Rootly
2w
Save
Mark Applied
Hide
Staff Network Reliability Engineer (Cloud Operations)
Mountain View, California, United States
HybridFull Time
Skylo
Skylo: Provides direct-to-device satellite connectivity for mobile and IoT devices.
8+ YOE8+ years cloud/infrastructure/SRE experience with Kubernetes, hybrid cloud operations, observability, database and storage reliability, GitOps, and on-call ownership in 24x7 environments.
Kubernetes, GKE, GCP, kubectl, Pub/Sub, Cloud SQL, Prometheus, VictoriaMetrics, Grafana, OpenTelemetry, PostgreSQL, Redis, ArgoCD, Helm, Terraform, Ansible, Ceph, Rook, Harvester, KubeVirt, KVM, Loki, ELK, Flux CD, Go, Python, BGP, VXLAN, EVPN
1w
Save
Mark Applied
Hide
Senior Site Reliability Engineer Platform Cloud Foundations Engineer
San Jose, California, United States
$64k-$130k/yr OnsiteFull Time
Tata Consultancy Services
Tata Consultancy ServicesNational Stock Exchange of India: TCS: Global provider of IT services, consulting, and business solutions.
8+ YOE8+ years SRE/platform engineering experience with AWS multi-account, Terraform, automation (Python/Go/Ruby), cloud governance, and strong documentation and communication skills.
AWS Organizations, IAM, Terraform, Python, Go, Ruby, Control Tower, Account Factory for Terraform, CloudFormation, EventBridge, Lambda, SQS, IAM Identity Center, GCP
5d
Save
Mark Applied
Hide
Senior Site Reliability Engineer - Cloud
Santa Clara, California, United States
$168k-$265k/yr OnsiteFull Time
NVIDIA
NVIDIANASDAQ: NVDA: Designs graphics processing units and artificial intelligence hardware.
8+ YOEMS/BS or equivalent experience, 8+ years supporting live-site production, SRE on-call experience, strong Kubernetes and Python skills, Akamai/CDN and AWS experience, incident management and automation focus.
Akamai Edge Redirector Cloudlets, Akamai Forward Rewrite Cloudlets, Akamai Cloudlets Policy Manager, Akamai CDN, WAF, AWS, Kubernetes, Python
2w
Save
Mark Applied
Hide
Cloud Engineer
Palo Alto, California, United States
$125k-$200k/yr OnsiteFull Time
Sage Care
Sage Care: AI-powered care navigation and scheduling platform for health systems.
4+ YOE4+ years DevOps/SRE experience with GCP, Kubernetes (GKE), Terraform, CI/CD, and Bazel; strong networking, IAM, cloud security, and production reliability skills.
GCP, Terraform, Kubernetes, GKE, Bazel, CI/CD
1mo
Save
Mark Applied
Hide
Site Reliability Engineer
Santa Clara or St. Louis or Bangalore or London or Paris or Melbourne or Taipei or Tokyo
OnsiteFull Time
Netskope
NetskopeNASDAQ: NTSK: Cloud-native cybersecurity and data protection platform for enterprises.
3+ YOEBachelor's in CS/Engineering or equivalent; 3+ years building/managing complex systems (including 1-2 years SRE); experience with cloud services, microservices, availability/performance optimization, debugging, and strong communication.
Python, C, C++, Go, Rust, Docker, Kubernetes, AWS, GCP, KVM, OpenNebula, OpenStack, TCP/IP
1mo
Save
Mark Applied
Hide
Principal Site Reliability Engineer, Google Cloud
Atlanta or Milpitas
$240k-$250k/yr HybridFull Time
Saviynt
Saviynt: Provides AI-powered identity governance and cloud security platforms.
9+ YOE9+ years in platform/infra/SRE roles, deep Kubernetes and GCP expertise, strong Go and Python skills, experience with CI/CD, event-driven systems, observability, distributed systems, and building shared platform services.
Go (Golang), Python, Kubernetes, GCP, AWS, Azure, Kafka, RMQ, NATS, Google Pub/Sub, GitLab CI, ArgoCD, Prometheus, Grafana, ELK stack, Datadog, Envoy, Istio, MySQL, PostgresSQL
1mo
Save
Mark Applied
Hide
Senior Systems Reliability Engineer
Pune or San Jose or Durham or Mexico City or Bangalore or Hoofddorp or Belgrade or Barcelona or Singapore or Sydney or Tokyo
HybridFull Time
Nutanix
NutanixNASDAQ: NTNX: Sells cloud software and hyperconverged infrastructure for enterprises.
7+ YOE7+ years SRE experience with networking, virtualization (VMware ESXi), Linux, cloud and strong customer-facing troubleshooting and communication skills.
VMware ESXi, VMware, Linux, DevOps, Cloud, Citrix, Microsoft
3w
Save
Mark Applied
Hide
Site Reliability Engineer
Santa Clara, California, United States
$230k-$250k/yr OnsiteFull Time
Forward Networks
Forward Networks: Provides a digital twin platform for enterprise network management.
6+ YOE6+ years SRE/DevOps experience in SaaS/cloud, strong networking fundamentals, Kubernetes, observability (Prometheus/Grafana/Datadog/Splunk), Python/Bash automation, cloud and IaC (AWS/GCP/Azure, Terraform/Ansible), and incident response ownership.
Kubernetes, Prometheus, Grafana, Datadog, Splunk, Python, Bash, AWS, GCP, Azure, Terraform, Ansible
3w
Save
Mark Applied
Hide
Senior Site Reliability Engineer
Sunnyvale, California, United States
$90k-$180k/yr OnsiteFull Time
Abbott
AbbottNYSE: ABT: Manufactures medical devices, diagnostics, and nutritional health products.
Ensure reliability, scalability, and performance of a medical-device remote monitoring platform; expertise in cloud (Azure), Kubernetes, observability, automation, and incident management; bachelor's in a technical discipline.
Python, Go, Bash, PowerShell, Microsoft Azure, Azure Kubernetes Service (AKS), Azure Monitor, Azure DevOps, Azure Policy, Kubernetes, Docker, Prometheus, Grafana, ELK/EFK, Datadog, Linux
3w
Save
Mark Applied
Hide
Senior Site Reliability Engineer
San Mateo, California, United States
$130k-$200k/yr OnsiteFull Time
IXL Learning
IXL Learning: Provides personalized digital learning platforms and educational resources.
6+ YOEBachelor's degree,6+ years SRE/software engineering,experience with OO and scripting languages,cloud (AWS/GCP),Docker/Kubernetes,monitoring,on-call availability,strong troubleshooting and communication skills.
Java, C++, C, Python, Bash, Perl, AWS, GCP, Docker, Kubernetes
4w
Save
Mark Applied
Hide
Senior Site Reliability Engineer
Palo Alto or Pittsburgh
$179k-$269k/yr OnsiteFull Time
Latitude AI
Latitude AI: Developing automated driving technology for next-generation Ford vehicles.
4+ YOEBachelor's degree in engineering/computer science (or higher) with 4+ years experience (or equivalent), strong Linux, networking, Go/Python development, cloud (AWS/GCP), Kubernetes, IaC, monitoring and SLO experience.
Go, Python, AWS, GCP, Terraform, CloudFormation, Kubernetes, Prometheus, Elasticsearch, Loki, Jaeger, Tempo, Linux
1w
Save
Mark Applied
Hide
Senior Site Reliability Engineer - Core Cloud Platform
San Francisco or San Jose or Bellevue
$240k-$356k/yr HybridFull Time
Lambda
Lambda: Provides high-performance GPU cloud infrastructure for AI development.
7+ YOE7+ years SRE or production infrastructure experience, deep Kubernetes and Terraform knowledge, experience with observability and SLOs, proficiency in Go or Python, on-call and incident leadership experience.
Kubernetes, Terraform, Argo CD, Flux, Helm, Kustomize, OpenTelemetry, Prometheus, Grafana, Datadog, Go, Python, etcd, GitOps
2mo
Save
Mark Applied
Hide
Senior Site Reliability Engineer
Palo Alto, California, United States
$200k-$400k/yr HybridFull Time
Nectar Social
Nectar Social: AI platform for social commerce and community management.
5+ YOE5+ years operating production systems; cloud (AWS); infrastructure as code; programming; startup environment; reliability-focused with cost awareness.
AWS, Pulumi, Postgres, ClickHouse, Turbopuffer, Temporal
3w
Save
Mark Applied
Hide
Senior Site Reliability Engineer
Sunnyvale or Sylmar
$90k-$180k/yr OnsiteFull Time
Abbott
AbbottNYSE: ABT: Provides medical devices, diagnostics, and science-based nutritional products.
Senior SRE with strong distributed systems, cloud (Azure), Kubernetes, observability, automation, incident management, and cross-functional communication skills for a medical device remote monitoring platform.
Python, Go, Bash, PowerShell, Microsoft Azure, Azure Kubernetes Service (AKS), Azure Monitor, Azure DevOps, Azure Policy, Kubernetes, Docker, Prometheus, Grafana, ELK, EFK, Datadog, Linux
1w
Save
Mark Applied
Hide
Sr. Site Reliability Engineer (Starlink)
Hawthorne or Palo Alto or Redmond
$165k-$270k/yr OnsiteFull Time
SpaceX
SpaceX: Designs and launches advanced rockets and satellite internet constellations.
5+ YOEBachelor's in CS/engineering/math with 5 years software experience or 7+ years SRE/DevOps experience; Linux experience required; Kubernetes, Kafka, cloud-native tooling, and programming in Python/Go/Java/C#/Scala preferred.
Linux, Kubernetes, Istio, Apache Kafka, Apache Spark, HBase, HDFS, Apache Flink, Python, C#, Java, Scala, Go
5d
Save
Mark Applied
Hide
Sr. Database Reliability Engineer
San Jose, California, United States
$139k-$258k/yr OnsiteFull Time
Adobe
AdobeNASDAQ: ADBE: Provides software for digital media creation and marketing analytics
7+ YOE7+ years operating highly available database platforms; strong experience with MongoDB/Cassandra/MySQL/PostgreSQL, cloud (AWS/Azure), managed DB services, IaC (Terraform/Chef/Ansible), Kubernetes/Docker, Python; bachelor's or equivalent experience.
MongoDB, Cassandra, MySQL, PostgreSQL, Percona XtraDB Cluster, MariaDB Galera Cluster, AWS, Azure, Amazon RDS, Keyspaces, DynamoDB, Azure SQL, Cosmos DB, MongoDB Atlas, Terraform, Chef, Ansible, Kubernetes, Docker, Python
2mo
Save
Mark Applied
Hide
Principal Site Reliability Engineer
Santa Clara, California, United States
$152k-$245k/yr OnsiteFull Time
Palo Alto Networks
Palo Alto NetworksNASDAQ: PANW: Provides enterprise-grade network, cloud, and endpoint security software.
BS or MS in CS or related field; expertise in configuration management (Ansible, Terraform, Kubernetes); Python and/or Go; Kubernetes with autoscaling; production engineering/DevOps/SRE experience; public cloud (GCP/AWS); Linux networking; CI/CD with GitLab/GitHub; distributed systems; strong communication; ownership and monitoring as code.
Kubernetes, Docker, GCP, AWS, Ansible, Terraform, Vault, GitLab, Spinnaker, Pub/Sub, Bigtable, Memorystore, BigQuery, RabbitMQ, Kafka, MySQL, Python, Go, Shell scripting, Golang
3w
Save
Mark Applied
Hide
Staff Site Reliability Engineer, Quota SRE
Sunnyvale, California, United States
$207k-$301k/yr OnsiteFull Time
Google
GoogleNASDAQ: GOOGL: Provides online search, advertising, cloud computing, and consumer electronics.
8+ YOEBachelor's degree or equivalent,8 years software/systems engineering experience,5 years SRE experience,5 years software design experience,EMR not mentioned; strong troubleshooting and stakeholder management skills.
Google Cloud, Quotaserver, Bouncer, Slicer