160 cloud reliability engineer jobs at 112 companies in Benicia, CA

2mo
Save
Mark Applied
Hide
Cloud Site Reliability Engineer
San Jose or Palo Alto
HybridFull Time
SambaNova Systems
SambaNova Systems: Develops custom AI hardware and software for enterprise computing.
3+ YOE3-5+ years SRE/DevOps in public cloud; Bachelor's degree or equivalent; Python/Go/Java; Docker/Kubernetes; monitoring/observability tools; IaC; CI/CD.
Docker, Kubernetes, Prometheus, Grafana, ELK Stack, Datadog, Terraform, CloudFormation, Jenkins, GitHub Actions, ArgoCD, Python, Go, Java
4w
Save
Mark Applied
Hide
Senior Reliability Engineer, DGX Cloud
Santa Clara or United States
$168k-$334k/yr HybridFull Time
NVIDIA
NVIDIANASDAQ: NVDA: Designs GPU-accelerated computing and artificial intelligence hardware.
10+ YOE10+ years running large-scale production systems, strong software engineering (Go/Python), SLO program experience, incident response leadership, chaos engineering and failure-injection expertise, ability to influence across teams.
Go, Python, Prometheus, OpenTelemetry, Grafana, PagerDuty, Rootly
4w
Save
Mark Applied
Hide
Senior Reliability Engineer, DGX Cloud
Santa Clara or United States
$168k-$334k/yr HybridFull Time
NVIDIA
NVIDIANASDAQ: NVDA: Designs graphics processing units and artificial intelligence hardware.
10+ YOE10+ years running large-scale production systems; strong software engineering in Go or Python; SLO program experience; chaos engineering and failure-injection experience; ability to lead incident response and influence cross-team.
Go, Python, Prometheus, OpenTelemetry, Grafana, PagerDuty, Rootly
3w
Save
Mark Applied
Hide
Staff Site Reliability Engineer, Cloud Reliability Intelligence
Sunnyvale, California, United States
$207k-$301k/yr OnsiteFull Time
Google
GoogleNASDAQ: GOOGL: Provides online search, advertising, cloud computing, and consumer electronics.
8+ YOEBachelor's degree or equivalent, 8+ years experience with data structures and algorithms, 3+ years leading distributed systems projects and in technical leadership; experience with full-stack architectures and LLM/Generative AI preferred.
Go, Java, TypeScript, Angular, LLMs, Generative AI
2mo
Save
Mark Applied
Hide
Site Reliability Engineer (SRE)
San Francisco, California, United States
$350k-$475k/yr OnsiteFull Time
Thinking Machines Lab
Thinking Machines Lab: Builds advanced multimodal AI models and model optimization infrastructure.
Bachelor's degree or equivalent experience; distributed systems, cloud infrastructure or SRE; reliability tooling; incident response; strong cross-team communication.
Kubernetes, Docker, Cloud Platforms, Monitoring Tools, Automation
2mo
Save
Mark Applied
Hide
Senior Site Reliability Engineer
San Francisco, California, United States
$167k-$226k/yr HybridFull Time
Drata
Drata: Automated security and compliance platform for businesses.
6+ YOE6+ years of Site Reliability Engineering or related cloud engineering experience; strong cloud, Terraform, Docker, Linux; Datadog monitoring; CI/CD with GitHub Actions; incident management; AI experience preferred.
Terraform, Docker, Git, Linux, Datadog, GitHub Actions, Python, Bash, AWS, ECS, Kubernetes, MySQL
1mo
Save
Mark Applied
Hide
Senior Site Reliability Engineer
Oakland, California, United States
$175k-$210k/yr HybridFull Time
Fivetran
Fivetran: Automates data movement into cloud data warehouses.
5+ YOE5+ years SaaS experience; managed Kubernetes, cloud platforms (AWS/GCP/Azure), Terraform/Ansible/ArgoCD; Python/Shell scripting, Linux admin, PostgreSQL; incident response and reliability engineering experience.
Kubernetes, EKS, AKS, GKE, PostgreSQL, ArgoCD, Terraform, Ansible, Python, Shell, Go, Java, AWS, GCP, Azure, Grafana, Buildkite, Temporal, Pulumi, Linux, VPN, PrivateLink, Private Service Connect (GCP)
6d
Save
Mark Applied
Hide
Site Reliability Engineer
San Francisco, California, United States
HybridFull Time
Runloop
Runloop: Provides infrastructure and secure sandboxes for AI agents.
5+ YOE5+ years software engineering experience with 3+ years in SRE/DevOps, strong Python or Go skills, containerization, cloud infra, monitoring, networking, Linux administration, on‑call and incident management.
AWS, GCP, Azure, Grafana, Prometheus, Datadog, Python, Go, Docker, Kubernetes, Terraform, Pulumi, Sentry, RUM, CI/CD
2w
Save
Mark Applied
Hide
Site Reliability Engineer
Santa Clara or St. Louis or Bangalore or London or Paris or Melbourne or Taipei or Tokyo
OnsiteFull Time
Netskope
NetskopeNASDAQ: NTSK: Cloud-native cybersecurity and data protection platform for enterprises.
3+ YOEBachelor's in CS/Engineering or equivalent; 3+ years building/managing complex systems (including 1-2 years SRE); experience with cloud services, microservices, availability/performance optimization, debugging, and strong communication.
Python, C, C++, Go, Rust, Docker, Kubernetes, AWS, GCP, KVM, OpenNebula, OpenStack, TCP/IP
2w
Save
Mark Applied
Hide
Principal Site Reliability Engineer, Google Cloud
Atlanta or Milpitas
$240k-$250k/yr HybridFull Time
Saviynt
Saviynt: Provides AI-powered identity governance and cloud security platforms.
9+ YOE9+ years in platform/infra/SRE roles, deep Kubernetes and GCP expertise, strong Go and Python skills, experience with CI/CD, event-driven systems, observability, distributed systems, and building shared platform services.
Go (Golang), Python, Kubernetes, GCP, AWS, Azure, Kafka, RMQ, NATS, Google Pub/Sub, GitLab CI, ArgoCD, Prometheus, Grafana, ELK stack, Datadog, Envoy, Istio, MySQL, PostgresSQL
5d
Save
Mark Applied
Hide
Site Reliability Engineer
Santa Clara, California, United States
$230k-$250k/yr OnsiteFull Time
Forward Networks
Forward Networks: Provides a digital twin platform for enterprise network management.
6+ YOE6+ years SRE/DevOps experience in SaaS/cloud, strong networking fundamentals, Kubernetes, observability (Prometheus/Grafana/Datadog/Splunk), Python/Bash automation, cloud and IaC (AWS/GCP/Azure, Terraform/Ansible), and incident response ownership.
Kubernetes, Prometheus, Grafana, Datadog, Splunk, Python, Bash, AWS, GCP, Azure, Terraform, Ansible
3d
Save
Mark Applied
Hide
Senior Site Reliability Engineer
Sunnyvale, California, United States
$90k-$180k/yr OnsiteFull Time
Abbott
AbbottNYSE: ABT: Manufactures medical devices, diagnostics, and nutritional health products.
Ensure reliability, scalability, and performance of a medical-device remote monitoring platform; expertise in cloud (Azure), Kubernetes, observability, automation, and incident management; bachelor's in a technical discipline.
Python, Go, Bash, PowerShell, Microsoft Azure, Azure Kubernetes Service (AKS), Azure Monitor, Azure DevOps, Azure Policy, Kubernetes, Docker, Prometheus, Grafana, ELK/EFK, Datadog, Linux
4d
Save
Mark Applied
Hide
Senior Site Reliability Engineer
San Mateo, California, United States
$130k-$200k/yr OnsiteFull Time
IXL Learning
IXL Learning: Provides personalized digital learning platforms and educational resources.
6+ YOEBachelor's degree,6+ years SRE/software engineering,experience with OO and scripting languages,cloud (AWS/GCP),Docker/Kubernetes,monitoring,on-call availability,strong troubleshooting and communication skills.
Java, C++, C, Python, Bash, Perl, AWS, GCP, Docker, Kubernetes
1w
Save
Mark Applied
Hide
Senior Site Reliability Engineer
Palo Alto or Pittsburgh
$179k-$269k/yr OnsiteFull Time
Latitude AI
Latitude AI: Developing automated driving technology for next-generation Ford vehicles.
4+ YOEBachelor's degree in engineering/computer science (or higher) with 4+ years experience (or equivalent), strong Linux, networking, Go/Python development, cloud (AWS/GCP), Kubernetes, IaC, monitoring and SLO experience.
Go, Python, AWS, GCP, Terraform, CloudFormation, Kubernetes, Prometheus, Elasticsearch, Loki, Jaeger, Tempo, Linux
1w
Save
Mark Applied
Hide
Senior Database Reliability Engineer
Southfield or San Francisco or Seattle or Boston or New York City or Los Angeles or San Diego or United States
$104k-$153k/yr RemoteFull Time
Credit Acceptance
Credit AcceptanceNASDAQ: CACC: Provides vehicle financing programs for subprime credit consumers.
7+ YOE7+ years in DB engineering/DBA/DBRE roles; strong SQL, performance tuning, HA, backup/recovery; bachelor's or equivalent experience; experience with cloud DBs, automation, observability, and on-call rotations.
Terraform, Ansible, AWS DMS, AWS RDS / Aurora (PostgreSQL preferred), DynamoDB, Oracle, SQL Server, MongoDB, MySQL, Datadog, Grafana, Prometheus, OpenTelemetry, Jenkins, GitHub Actions, Python, CyberArk, Liquibase, Flyway, Hibernate, JPA, SQLAlchemy, REST APIs, Linux/Unix, SQL
2w
Save
Mark Applied
Hide
Site Reliability Engineer
San Francisco, California, United States
OnsiteFull Time
Specter: Building a software-defined perception engine for the physical world.
Strong Linux administration, experience with edge/on‑prem hardware and cloud (AWS), networking fundamentals, scripting in Python/Go/Bash, containerization (Docker, Kubernetes) and embedded/firmware familiarity; on‑call participation.
AWS, Bash, C, Docker, Go, Kubernetes, Linux, Python, Rust, SSH, DNS, VPN, IAM
2w
Save
Mark Applied
Hide
Senior Site Reliability Engineer, Robotics & Cloud Infrastructure
Brooklyn or New York City or Richmond or Europe
$164k-$220k/yr RemoteFull Time
Bedrock Ocean Exploration
Bedrock Ocean Exploration: Maps the ocean floor using autonomous underwater robotic vehicles.
5+ YOE5+ years SRE/DevOps experience with on-call ownership; strong automation using Python/Go/Bash; Terraform and AWS hands-on; containerization (Docker, Kubernetes); observability (Prometheus, Grafana); Linux and networking expertise; East Coast location and US work authorization required.
Python, Go, Bash, Terraform, AWS, Docker, Kubernetes, Prometheus, Grafana, ROS 2, ROS, Jetson, Linux, IAM
1mo
Save
Mark Applied
Hide
Site Reliability Engineer
United States or San Francisco or New York City
$101k-$199k/yr OnsiteFull Time
Microsoft
MicrosoftNASDAQ: MSFT: Develops software, services, devices, and cloud computing solutions.
1+ YOEMaster's or Bachelor's in CS/IT (or equivalent experience), 1+ years managing physical infrastructure, on-call experience, experience with large-scale cloud/distributed systems preferred, and ability to pass Microsoft security screening.
Azure, InfiniBand, GPUs
3w
Save
Mark Applied
Hide
Customer Reliability Engineer, Airflow
United States or San Francisco or Boston or Atlanta or Austin or Washington D.C. or Raleigh or Pittsburgh or Philadelphia or New York City or Miami or Columbus
$125k-$130k/yr RemoteFull Time
Astronomer
Astronomer: Managed data orchestration platform powered by Apache Airflow.
4+ YOEData engineering background, 4 years Python, 1 year Airflow administration/DAG creation, Kubernetes/Docker experience, cloud provider (AWS/GCP/Azure) experience, troubleshooting, strong communication, and mentoring experience.
Apache Airflow, Python, Kubernetes, Docker, AWS, GCP, Azure, SQL, PostgreSQL, Databricks, Snowflake, Redshift, dbt, Zoom
5d
Save
Mark Applied
Hide
Site Reliability Engineer
Research Triangle Park or San Jose or Milpitas or Richardson or Santa Clara
$127k-$182k/yr HybridFull Time
Cisco
CiscoNASDAQ: CSCO: Develops and sells networking hardware and cybersecurity software.
5+ YOE5+ years SRE/Cloud Ops experience, Docker and Kubernetes proficiency, scripting in Python/Go/Bash, monitoring and incident response experience, Linux and networking knowledge, CI/CD and IaC familiarity, bachelor’s degree or equivalent.
Docker, Kubernetes, Python, Go, Bash, Git