68 cloud reliability engineer jobs at 26 companies in Hollister, CA

3w
Save
Mark Applied
Hide
Cloud Site Reliability Engineer - DCS Cloud
San Jose, California, United States
OnsiteFull Time
ByteDance
ByteDance: Developing AI-driven content platforms and mobile applications.
2+ YOEBachelor's degree in CS or related,2+ years in Linux operations/SRE/DevOps,programming in Go/Python/C++,cloud and reliability practices experience,strong troubleshooting and communication skills.
Go, Python, C++, Linux, OCI, AWS, Azure, GCP, KVM, QEMU, Docker, Kubernetes, containerd, cgroups, namespaces, CUDA, MIG
1mo
Save
Mark Applied
Hide
Senior Reliability Engineer, DGX Cloud
Santa Clara or United States
$168k-$334k/yr HybridFull Time
NVIDIA
NVIDIANASDAQ: NVDA: Designs GPU-accelerated computing and artificial intelligence hardware.
10+ YOE10+ years running large-scale production systems, strong software engineering (Go/Python), SLO program experience, incident response leadership, chaos engineering and failure-injection expertise, ability to influence across teams.
Go, Python, Prometheus, OpenTelemetry, Grafana, PagerDuty, Rootly
5d
Save
Mark Applied
Hide
Staff Reliability Engineer
Santa Clara, California, United States
$167k-$291k/yr RemoteFull Time
ServiceNow
ServiceNowNYSE: NOW: Provides a cloud platform for automating enterprise digital workflows.
8+ YOE8+ years SRE/Platform/DevOps experience, strong Kubernetes and cloud-native platform skills, automation and CI/CD expertise, software engineering with Python/Go/Java/Ruby, observability and reliability knowledge.
Kubernetes, Python, Go, Java, Ruby, GitLab CI/CD, Argo CD, Flux, Playwright, Selenium, Cypress, REST Assured, PyTest, JUnit, TestNG, Ansible, Terraform, Helm, Argo Workflows, Kustomize, Istio, Linkerd, Gateway API, Ingress, Prometheus, OpenTelemetry, AWS (EKS), Azure (AKS), Google Cloud (GKE), GitOps
1w
Save
Mark Applied
Hide
Senior Site Reliability Engineer Platform Cloud Foundations Engineer
San Jose, California, United States
$64k-$130k/yr OnsiteFull Time
Tata Consultancy Services
Tata Consultancy ServicesNational Stock Exchange of India: TCS: Global provider of IT services, consulting, and business solutions.
8+ YOE8+ years SRE/platform engineering experience with AWS multi-account, Terraform, automation (Python/Go/Ruby), cloud governance, and strong documentation and communication skills.
AWS Organizations, IAM, Terraform, Python, Go, Ruby, Control Tower, Account Factory for Terraform, CloudFormation, EventBridge, Lambda, SQS, IAM Identity Center, GCP
3d
Save
Mark Applied
Hide
Senior Site Reliability Engineer - Cloud
Santa Clara, California, United States
$168k-$265k/yr OnsiteFull Time
NVIDIA
NVIDIANASDAQ: NVDA: Designs graphics processing units and artificial intelligence hardware.
8+ YOEMS/BS or equivalent experience, 8+ years supporting live-site production, SRE on-call experience, strong Kubernetes and Python skills, Akamai/CDN and AWS experience, incident management and automation focus.
Akamai Edge Redirector Cloudlets, Akamai Forward Rewrite Cloudlets, Akamai Cloudlets Policy Manager, Akamai CDN, WAF, AWS, Kubernetes, Python
1mo
Save
Mark Applied
Hide
Site Reliability Engineer
Santa Clara or St. Louis or Bangalore or London or Paris or Melbourne or Taipei or Tokyo
OnsiteFull Time
Netskope
NetskopeNASDAQ: NTSK: Cloud-native cybersecurity and data protection platform for enterprises.
3+ YOEBachelor's in CS/Engineering or equivalent; 3+ years building/managing complex systems (including 1-2 years SRE); experience with cloud services, microservices, availability/performance optimization, debugging, and strong communication.
Python, C, C++, Go, Rust, Docker, Kubernetes, AWS, GCP, KVM, OpenNebula, OpenStack, TCP/IP
1mo
Save
Mark Applied
Hide
Principal Site Reliability Engineer, Google Cloud
Atlanta or Milpitas
$240k-$250k/yr HybridFull Time
Saviynt
Saviynt: Provides AI-powered identity governance and cloud security platforms.
9+ YOE9+ years in platform/infra/SRE roles, deep Kubernetes and GCP expertise, strong Go and Python skills, experience with CI/CD, event-driven systems, observability, distributed systems, and building shared platform services.
Go (Golang), Python, Kubernetes, GCP, AWS, Azure, Kafka, RMQ, NATS, Google Pub/Sub, GitLab CI, ArgoCD, Prometheus, Grafana, ELK stack, Datadog, Envoy, Istio, MySQL, PostgresSQL
4w
Save
Mark Applied
Hide
Senior Systems Reliability Engineer
Pune or San Jose or Durham or Mexico City or Bangalore or Hoofddorp or Belgrade or Barcelona or Singapore or Sydney or Tokyo
HybridFull Time
Nutanix
NutanixNASDAQ: NTNX: Sells cloud software and hyperconverged infrastructure for enterprises.
7+ YOE7+ years SRE experience with networking, virtualization (VMware ESXi), Linux, cloud and strong customer-facing troubleshooting and communication skills.
VMware ESXi, VMware, Linux, DevOps, Cloud, Citrix, Microsoft
3w
Save
Mark Applied
Hide
Site Reliability Engineer
Santa Clara, California, United States
$230k-$250k/yr OnsiteFull Time
Forward Networks
Forward Networks: Provides a digital twin platform for enterprise network management.
6+ YOE6+ years SRE/DevOps experience in SaaS/cloud, strong networking fundamentals, Kubernetes, observability (Prometheus/Grafana/Datadog/Splunk), Python/Bash automation, cloud and IaC (AWS/GCP/Azure, Terraform/Ansible), and incident response ownership.
Kubernetes, Prometheus, Grafana, Datadog, Splunk, Python, Bash, AWS, GCP, Azure, Terraform, Ansible
2w
Save
Mark Applied
Hide
Senior Site Reliability Engineer
Sunnyvale, California, United States
$90k-$180k/yr OnsiteFull Time
Abbott
AbbottNYSE: ABT: Manufactures medical devices, diagnostics, and nutritional health products.
Ensure reliability, scalability, and performance of a medical-device remote monitoring platform; expertise in cloud (Azure), Kubernetes, observability, automation, and incident management; bachelor's in a technical discipline.
Python, Go, Bash, PowerShell, Microsoft Azure, Azure Kubernetes Service (AKS), Azure Monitor, Azure DevOps, Azure Policy, Kubernetes, Docker, Prometheus, Grafana, ELK/EFK, Datadog, Linux
1w
Save
Mark Applied
Hide
Senior Site Reliability Engineer - Core Cloud Platform
San Francisco or San Jose or Bellevue
$240k-$356k/yr HybridFull Time
Lambda
Lambda: Provides high-performance GPU cloud infrastructure for AI development.
7+ YOE7+ years SRE or production infrastructure experience, deep Kubernetes and Terraform knowledge, experience with observability and SLOs, proficiency in Go or Python, on-call and incident leadership experience.
Kubernetes, Terraform, Argo CD, Flux, Helm, Kustomize, OpenTelemetry, Prometheus, Grafana, Datadog, Go, Python, etcd, GitOps
2w
Save
Mark Applied
Hide
Senior Site Reliability Engineer
Sunnyvale or Sylmar
$90k-$180k/yr OnsiteFull Time
Abbott
AbbottNYSE: ABT: Provides medical devices, diagnostics, and science-based nutritional products.
Senior SRE with strong distributed systems, cloud (Azure), Kubernetes, observability, automation, incident management, and cross-functional communication skills for a medical device remote monitoring platform.
Python, Go, Bash, PowerShell, Microsoft Azure, Azure Kubernetes Service (AKS), Azure Monitor, Azure DevOps, Azure Policy, Kubernetes, Docker, Prometheus, Grafana, ELK, EFK, Datadog, Linux
3d
Save
Mark Applied
Hide
Sr. Database Reliability Engineer
San Jose, California, United States
$139k-$258k/yr OnsiteFull Time
Adobe
AdobeNASDAQ: ADBE: Provides software for digital media creation and marketing analytics
7+ YOE7+ years operating highly available database platforms; strong experience with MongoDB/Cassandra/MySQL/PostgreSQL, cloud (AWS/Azure), managed DB services, IaC (Terraform/Chef/Ansible), Kubernetes/Docker, Python; bachelor's or equivalent experience.
MongoDB, Cassandra, MySQL, PostgreSQL, Percona XtraDB Cluster, MariaDB Galera Cluster, AWS, Azure, Amazon RDS, Keyspaces, DynamoDB, Azure SQL, Cosmos DB, MongoDB Atlas, Terraform, Chef, Ansible, Kubernetes, Docker, Python
2mo
Save
Mark Applied
Hide
Principal Site Reliability Engineer
Santa Clara, California, United States
$152k-$245k/yr OnsiteFull Time
Palo Alto Networks
Palo Alto NetworksNASDAQ: PANW: Provides enterprise-grade network, cloud, and endpoint security software.
BS or MS in CS or related field; expertise in configuration management (Ansible, Terraform, Kubernetes); Python and/or Go; Kubernetes with autoscaling; production engineering/DevOps/SRE experience; public cloud (GCP/AWS); Linux networking; CI/CD with GitLab/GitHub; distributed systems; strong communication; ownership and monitoring as code.
Kubernetes, Docker, GCP, AWS, Ansible, Terraform, Vault, GitLab, Spinnaker, Pub/Sub, Bigtable, Memorystore, BigQuery, RabbitMQ, Kafka, MySQL, Python, Go, Shell scripting, Golang
2w
Save
Mark Applied
Hide
Staff Site Reliability Engineer, Quota SRE
Sunnyvale, California, United States
$207k-$301k/yr OnsiteFull Time
Google
GoogleNASDAQ: GOOGL: Provides online search, advertising, cloud computing, and consumer electronics.
8+ YOEBachelor's degree or equivalent,8 years software/systems engineering experience,5 years SRE experience,5 years software design experience,EMR not mentioned; strong troubleshooting and stakeholder management skills.
Google Cloud, Quotaserver, Bouncer, Slicer
4w
Save
Mark Applied
Hide
Senior Site Reliability Engineer, Global E-Commerce
San Jose, California, United States
$213k-$388k/yr OnsiteFull Time
TikTok
TikTok: Global short-form video hosting and social media platform.
5+ YOEBachelor's or equivalent,5+ years SRE/infra experience,proficiency in Go/Python/Java,strong Linux,networking and distributed systems knowledge,cloud-native production experience.
Go, Python, Java, Linux
3w
Save
Mark Applied
Hide
Senior Site Reliability Engineer (SRE) – CloudVision as a Service (CVaaS)
Santa Clara, California, United States
$101k-$161k/yr RemoteFull Time
Arista Networks
Arista NetworksNYSE: ANET: Provides cloud networking solutions and high-speed multilayer Ethernet switches.
5+ YOEBS/MS or equivalent experience,5+ years software engineering, experience with distributed databases/SaaS deployments, proficiency in Python/Golang/Bash, Kubernetes and cloud platform experience preferred.
Golang, Python, Ansible, Pulumi, Bash, Kubernetes, GKE, GCP
1mo
Save
Mark Applied
Hide
Site Reliability Engineer - USDS (Multiple Positions)
San Jose, California, United States
$188k-$259k/yr OnsiteFull Time
TikTok USDS Joint Venture
TikTok USDS Joint Venture: Operates and secures TikTok services for U.S. users.
1+ YOERequires degree in CS/Engineering/Information Systems/Data Science/Mathematics plus relevant experience; experience with Linux administration, monitoring, troubleshooting, SDLC and cloud-native operations.
Linux
1mo
Save
Mark Applied
Hide
Senior Manager, Networ Reliability Engineering
Santa Clara or Seattle or United States
$133k-$306k/yr OnsiteFull Time
Oracle
OracleNYSE: ORCL: Provides cloud infrastructure and enterprise software for global businesses.
5+ YOE3+ Mgmt5+ years network reliability engineering, 3+ years engineering/operations management, strong cloud networking and distributed systems expertise, proven people leadership, excellent communication and organizational skills.
OCI
3w
Save
Mark Applied
Hide
Contract Lead, Site Reliability Engineering — AI Accelerator Infrastructure
Santa Clara, California, United States
$195k-$285k/yr HybridContract, Full Time
d-Matrix: Develops high-performance semiconductor chips for generative AI inference.
15+ YOE5+ MgmtBachelor's in CS/EE,15+ years SRE/infrastructure engineering,5+ years leading SRE teams,deep Linux,Terraform,Ansible,Kubernetes,Prometheus/Grafana/Datadog,Python or Go,cloud (AWS/Azure/GCP).
Prometheus, Grafana, Datadog, Terraform, Ansible, Kubernetes, Python, Go, AWS, Azure, GCP, Slurm, LSF, InfiniBand, RoCE, NVLink