72 cloud reliability engineer jobs at 57 companies in Santa Rosa, CA
3w
Save
Mark Applied
Hide
3w
Site Reliability Engineer
San Francisco, California, United States
HybridFull Time
Runloop: Provides infrastructure and secure sandboxes for AI agents.
5+ YOE5+ years software engineering experience with 3+ years in SRE/DevOps, strong Python or Go skills, containerization, cloud infra, monitoring, networking, Linux administration, on‑call and incident management.
Ivo: AI-powered contract review and intelligence platform for legal teams.
5+ YOEMinimum 5 years experience; own uptime and reliability, define SLIs/SLOs/SLAs, design failover and disaster recovery, implement security controls, lead incident response; familiarity with cloud and LLM-driven systems.
Specter: Building a software-defined perception engine for the physical world.
Strong Linux administration, experience with edge/on‑prem hardware and cloud (AWS), networking fundamentals, scripting in Python/Go/Bash, containerization (Docker, Kubernetes) and embedded/firmware familiarity; on‑call participation.
Senior Site Reliability Engineer, Robotics & Cloud Infrastructure
Brooklyn or New York City or Richmond or Europe
$164k-$220k/yrRemoteFull Time
Bedrock Ocean Exploration: Maps the ocean floor using autonomous underwater robotic vehicles.
5+ YOE5+ years SRE/DevOps experience with on-call ownership; strong automation using Python/Go/Bash; Terraform and AWS hands-on; containerization (Docker, Kubernetes); observability (Prometheus, Grafana); Linux and networking expertise; East Coast location and US work authorization required.
Senior Site Reliability Engineer - Core Cloud Platform
San Francisco or San Jose or Bellevue
$240k-$356k/yrHybridFull Time
Lambda: Provides high-performance GPU cloud infrastructure for AI development.
7+ YOE7+ years SRE or production infrastructure experience, deep Kubernetes and Terraform knowledge, experience with observability and SLOs, proficiency in Go or Python, on-call and incident leadership experience.
United States or Kansas or Washington or California or Texas or Illinois or North Carolina or Colorado or Massachusetts or Pennsylvania or Virginia or Oregon or Nevada or Hawaii or New York or Georgia or Ohio or Arizona or Seattle or San Francisco or New York City
$110k-$183k/yrRemoteFull Time
Veeam: Data resilience and security for hybrid cloud environments
3+ YOE3+ years in software engineering with 1+ year in SRE/Platform/DevOps, cloud experience (Azure or comparable), observability (Prometheus, Grafana, OpenTelemetry, ELK), IaC (Terraform/Terragrunt/Pulumi), Kubernetes, CI/CD tooling, programming in TypeScript/JS, Go, Java, or C#, and experience in compliance-oriented environments.
Claryo: AI-powered spatial software for optimizing warehouse operations
3+ YOE3+ years SRE/infrastructure experience, strong Linux and networking fundamentals, experience with Kubernetes, cloud platforms, observability tooling, and debugging distributed systems in production.
Coupa: Cloud-based platform for managing and optimizing business expenditures.
8+ YOE8+ years hands-on DBA experience, deep MySQL expertise, scripting (Bash/Python/Ruby), cloud (AWS/RDS/Aurora) and automation experience, monitoring and HA/DR skills, ability to lead architecture and mentor engineers.
uRun: Infrastructure cloud for interactive, stateful AI inference.
7+ YOE7+ years in site reliability or infrastructure engineering; strong SLOs, incident response, and observability; Kubernetes and cloud (AWS); software engineering fundamentals; first SRE at a company.
Graphon: Developing graph-native AI models for multimodal data reasoning.
Proficient in Bash and Python; experience with infrastructure-as-code, Docker, CI/CD, multi-cloud deployments, networking and identity access; comfortable managing production environments and using AI tools.
Bash, Python, Infrastructure-as-code, Docker, CI/CD, AI tools
Senior Site Reliability Engineer, Production Engineer - ThousandEyes
San Francisco or Seattle or Austin or New York City
$165k-$241k/yrHybridFull Time
CiscoNASDAQ: CSCO: Develops and sells networking hardware and cybersecurity software.
5+ YOE5+ years experience; proficiency in Python or Go; expertise with Kubernetes, cloud (AWS), Unix/Linux; strong SRE principles, incident response, and security-minded engineering.
Python, Go, Kubernetes, Service Mesh, Prometheus, OpenTelemetry, ArgoCD, CNCF, AWS, Unix, Linux
San Francisco or Boston or Washington D.C. or Raleigh or Pittsburgh or Philadelphia or New York City or Miami or Columbus or Austin or United States
$125k-$130k/yrRemoteFull Time
Astronomer: Managed data orchestration platform powered by Apache Airflow.
5+ YOE5+ years with large cloud infrastructures, 3+ years Kubernetes, production distributed systems on AWS/GCP/Azure, strong Linux, Python scripting, DevOps/CI/CD, observability/monitoring, and customer-facing troubleshooting.
AutodeskNASDAQ: ADSK: Developing software for architecture, engineering, and entertainment industries.
7+ YOEU.S. citizen required. 7+ years SRE/platform/cloud experience; B.S. in CS/Engineering or equivalent; experience with large-scale cloud production systems, SLOs/SLIs, observability, incident management, automation, and IaC. Programming in Python/Go/Java/PowerShell/Bash.
Senior Site Reliability Engineer, Platform Infrastructure (Foundations)
San Francisco or Palo Alto
OnsiteFull Time
Anyscale: Cloud platform for scaling distributed machine learning applications.
3+ YOE3+ years writing production code; experience with distributed systems, Kubernetes, cloud (AWS/Azure/GCP); proficiency in Go and Python; familiarity with observability (Prometheus, Grafana); on-call experience.
Brasilia or San Francisco or Buenos Aires or Bogota
RemoteFull Time
Qdrant: Open-source vector search engine for AI retrieval and infrastructure.
5+ YOE5+ years in SRE/platform engineering, strong Go or Python skills, production Kubernetes experience, cloud (AWS/GCP/Azure) knowledge, automation and reliability focus, comfortable on-call.
CrewAI: Platform for orchestrating collaborative multi-agent AI systems.
Experience building and operating production SaaS infrastructure: cloud, containers, CI/CD, observability, secrets, databases, and automation using Python/Ruby/Go/Bash.