160 cloud reliability engineer jobs at 112 companies in Benicia, CA
2mo
Save
Mark Applied
Hide
2mo
Cloud Site Reliability Engineer
San Jose or Palo Alto
HybridFull Time
SambaNova Systems: Develops custom AI hardware and software for enterprise computing.
3+ YOE3-5+ years SRE/DevOps in public cloud; Bachelor's degree or equivalent; Python/Go/Java; Docker/Kubernetes; monitoring/observability tools; IaC; CI/CD.
NVIDIANASDAQ: NVDA: Designs GPU-accelerated computing and artificial intelligence hardware.
10+ YOE10+ years running large-scale production systems, strong software engineering (Go/Python), SLO program experience, incident response leadership, chaos engineering and failure-injection expertise, ability to influence across teams.
NVIDIANASDAQ: NVDA: Designs graphics processing units and artificial intelligence hardware.
10+ YOE10+ years running large-scale production systems; strong software engineering in Go or Python; SLO program experience; chaos engineering and failure-injection experience; ability to lead incident response and influence cross-team.
8+ YOEBachelor's degree or equivalent, 8+ years experience with data structures and algorithms, 3+ years leading distributed systems projects and in technical leadership; experience with full-stack architectures and LLM/Generative AI preferred.
Go, Java, TypeScript, Angular, LLMs, Generative AI
Drata: Automated security and compliance platform for businesses.
6+ YOE6+ years of Site Reliability Engineering or related cloud engineering experience; strong cloud, Terraform, Docker, Linux; Datadog monitoring; CI/CD with GitHub Actions; incident management; AI experience preferred.
Runloop: Provides infrastructure and secure sandboxes for AI agents.
5+ YOE5+ years software engineering experience with 3+ years in SRE/DevOps, strong Python or Go skills, containerization, cloud infra, monitoring, networking, Linux administration, on‑call and incident management.
Santa Clara or St. Louis or Bangalore or London or Paris or Melbourne or Taipei or Tokyo
OnsiteFull Time
NetskopeNASDAQ: NTSK: Cloud-native cybersecurity and data protection platform for enterprises.
3+ YOEBachelor's in CS/Engineering or equivalent; 3+ years building/managing complex systems (including 1-2 years SRE); experience with cloud services, microservices, availability/performance optimization, debugging, and strong communication.
Saviynt: Provides AI-powered identity governance and cloud security platforms.
9+ YOE9+ years in platform/infra/SRE roles, deep Kubernetes and GCP expertise, strong Go and Python skills, experience with CI/CD, event-driven systems, observability, distributed systems, and building shared platform services.
AbbottNYSE: ABT: Manufactures medical devices, diagnostics, and nutritional health products.
Ensure reliability, scalability, and performance of a medical-device remote monitoring platform; expertise in cloud (Azure), Kubernetes, observability, automation, and incident management; bachelor's in a technical discipline.
Python, Go, Bash, PowerShell, Microsoft Azure, Azure Kubernetes Service (AKS), Azure Monitor, Azure DevOps, Azure Policy, Kubernetes, Docker, Prometheus, Grafana, ELK/EFK, Datadog, Linux
IXL Learning: Provides personalized digital learning platforms and educational resources.
6+ YOEBachelor's degree,6+ years SRE/software engineering,experience with OO and scripting languages,cloud (AWS/GCP),Docker/Kubernetes,monitoring,on-call availability,strong troubleshooting and communication skills.
7+ YOE7+ years in DB engineering/DBA/DBRE roles; strong SQL, performance tuning, HA, backup/recovery; bachelor's or equivalent experience; experience with cloud DBs, automation, observability, and on-call rotations.
Specter: Building a software-defined perception engine for the physical world.
Strong Linux administration, experience with edge/on‑prem hardware and cloud (AWS), networking fundamentals, scripting in Python/Go/Bash, containerization (Docker, Kubernetes) and embedded/firmware familiarity; on‑call participation.
Senior Site Reliability Engineer, Robotics & Cloud Infrastructure
Brooklyn or New York City or Richmond or Europe
$164k-$220k/yrRemoteFull Time
Bedrock Ocean Exploration: Maps the ocean floor using autonomous underwater robotic vehicles.
5+ YOE5+ years SRE/DevOps experience with on-call ownership; strong automation using Python/Go/Bash; Terraform and AWS hands-on; containerization (Docker, Kubernetes); observability (Prometheus, Grafana); Linux and networking expertise; East Coast location and US work authorization required.
MicrosoftNASDAQ: MSFT: Develops software, services, devices, and cloud computing solutions.
1+ YOEMaster's or Bachelor's in CS/IT (or equivalent experience), 1+ years managing physical infrastructure, on-call experience, experience with large-scale cloud/distributed systems preferred, and ability to pass Microsoft security screening.
United States or San Francisco or Boston or Atlanta or Austin or Washington D.C. or Raleigh or Pittsburgh or Philadelphia or New York City or Miami or Columbus
$125k-$130k/yrRemoteFull Time
Astronomer: Managed data orchestration platform powered by Apache Airflow.
4+ YOEData engineering background, 4 years Python, 1 year Airflow administration/DAG creation, Kubernetes/Docker experience, cloud provider (AWS/GCP/Azure) experience, troubleshooting, strong communication, and mentoring experience.
Research Triangle Park or San Jose or Milpitas or Richardson or Santa Clara
$127k-$182k/yrHybridFull Time
CiscoNASDAQ: CSCO: Develops and sells networking hardware and cybersecurity software.
5+ YOE5+ years SRE/Cloud Ops experience, Docker and Kubernetes proficiency, scripting in Python/Go/Bash, monitoring and incident response experience, Linux and networking knowledge, CI/CD and IaC familiarity, bachelor’s degree or equivalent.