48 sre manager jobs at 39 companies in Cotati, CA

4d
Save
Mark Applied
Hide
Principal Product Manager Lead
United States or Seattle or San Francisco or Sunnyvale or Raleigh or Boston or London or Lisbon or Bengaluru or Dublin or Kyiv or Chicago or New York City or Austin or Phoenix or Las Vegas or Washington or Dallas
$200k-$275k/yr OnsiteFull Time
SingleStore
SingleStore: Provides a distributed SQL database for real-time analytics.
Product leadership experience with distributed databases, query processing, AI agents or MCP servers, observability systems, SRE workflows, complex technical initiatives, and cross-functional people leadership.
Aura Analyst, AI Agents, MCP servers, Python UDFs, Cloud Functions, Container Services, SQrL, SRE, CPU, GPU
1mo
Save
Mark Applied
Hide
Senior SRE Engineer - San Francisco
San Francisco, California, United States
HybridFull Time
Plaud
Plaud: Develops AI-powered voice recorders and automated transcription software.
8+ YOE8+ years in SRE/Infrastructure/Platform engineering, strong cloud (AWS/GCP/Azure) and Kubernetes experience, on-call/incident management experience, proficiency in Go/Python/Java, and experience building observability and SLO-driven systems.
AWS, GCP, Azure, Kubernetes, Go, Python, Java, Cursor, GPT models, Gemini, Claude
3w
Save
Mark Applied
Hide
Site Reliability Manager, Play and Android SRE
San Francisco, California, United States
$262k-$365k/yr OnsiteFull Time
Google
GoogleNASDAQ: GOOGL: Provides online search, advertising, cloud computing, and consumer electronics.
8+ YOE8+ MgmtBachelor's in CS or equivalent,8+ years software engineering,8+ years people management,5+ years quality/reliability engineering,experience with distributed systems and cloud,strong leadership and partnership skills.
2mo
Save
Mark Applied
Hide
Head - SRE
Bengaluru or Las Vegas or San Francisco
HybridFull Time
Skillz
SkillzNYSE: SKLZ: Operates a platform for competitive multiplayer mobile gaming.
14+ YOE14+ years infrastructure engineering experience with public cloud (AWS), 5+ years running Kubernetes (EKS), leadership experience, observability/CI-CD expertise, cost optimization track record, and proficiency in Go, Python, or Java.
EKS, EC2, VPC, IAM, Cost Explorer, Savings Plans, Kubernetes, Istio, Datadog, Prometheus, Jaeger, X-Ray, ArgoCD, GitHub Actions, Go, Python, Java
2mo
Save
Mark Applied
Hide
Senior Product Manager
Toronto or San Francisco
HybridFull Time
PagerDuty
PagerDutyNYSE: PD: Platform for real-time incident response and digital operations automation.
5+ YOE5+ years product management; SaaS B2B; SRE/DevOps domain experience; workflow automation and integration tooling; security and permissions expertise; strong communication; leadership in roadmap execution.
Workflow automation, APIs, Telemetry analysis, Security and authorization models, SaaS platforms
3w
Save
Mark Applied
Hide
Staff Product Manager - Observability
Bellevue or San Francisco
$291k-$430k/yr HybridFull Time
Lambda
Lambda: Provides high-performance GPU cloud infrastructure for AI development.
7+ YOE7+ years product management experience including 3+ years on technical infrastructure or platform products; experience defining observability, SLIs/SLOs/SLA, partnering with SRE and fleet engineering; strong writing, influence, and execution skills.
NCCL, InfiniBand, PyTorch, DCGM, Datadog, Grafana, Prometheus
1mo
Save
Mark Applied
Hide
Staff Technical Program Manager
San Francisco, California, United States
$200k-$240k/yr HybridFull Time
Sentry
Sentry: Developer platform for error tracking and performance monitoring.
8+ YOE8+ years TPM experience in high-growth tech with platform/infrastructure/SRE/security exposure; proven cross-functional program leadership; incident management experience; strong analytical and communication skills.
4d
Save
Mark Applied
Hide
Principal Product Manager Lead
United States or Seattle or San Francisco or Sunnyvale or Raleigh or Boston or London or Lisbon or Bangalore or Dublin or Kyiv or Chicago or New York City or Austin or Phoenix or Las Vegas or Dallas
$200k-$275k/yr OnsiteFull Time
SingleStore
SingleStore: Real-time distributed SQL database for transactions and analytics.
Requires deep distributed database, query processing, performance optimization, observability, and AI product expertise, plus multi-team leadership, customer engagement, execution, communication, and mentorship skills.
Aura Analyst, AI Agents, MCP servers, Python UDFs, Cloud Functions, Container Services, SQrL, SRE
2mo
Save
Mark Applied
Hide
Engineering Manager, Cloud Platform
San Francisco, California, United States
$165k-$330k/yr HybridFull Time
Baseten
Baseten: Scalable infrastructure platform for deploying and serving AI models.
Experienced manager for cloud platform/SRE teams with strong Kubernetes and production infrastructure background, familiar with IaC and CI/CD tooling, recruiting and incident management skills.
Kubernetes, Terraform, CloudFormation, Pulumi, GitHub Actions, GitLab CI, CircleCI, Jenkins, Prometheus, ELK stack, Grafana, OpenTelemetry
2w
Save
Mark Applied
Hide
Engineering Manager, Infrastructure Engineering
Bellevue or Livingston or New York City or Sunnyvale or San Francisco
$182k-$242k/yr OnsiteFull Time
CoreWeave
CoreWeaveNASDAQ: CRWV: Cloud platform providing GPU-accelerated infrastructure for AI workloads.
7+ YOE3+ Mgmt3+ years engineering management and 7+ years technical experience in cloud operations/SRE; strong knowledge of cloud platforms, K8S, observability, incident management, and systems programming (Go); experience defining SLAs/SLOs and running on-call.
Prometheus, Grafana, Go, K8S, Redfish, BMC
1mo
Save
Mark Applied
Hide
Incident Manager
United States or Texas or San Francisco
$104k-$146k/yr RemoteFull Time
Databricks
Databricks: A unified platform for data analytics and artificial intelligence.
5+ YOE5+ years in incident management/SRE/production operations for cloud-native systems; lead high-severity incidents; cloud (AWS/Azure/GCP) and observability expertise; log analysis; scripting (Python/Go/Bash); BS/MS in CS/CE or related field.
AWS, Azure, GCP, Datadog, Elasticsearch, Splunk, Cloud Logging, OpenTelemetry, Prometheus, Grafana, Python, Go, Bash, Apache Spark, Delta Lake, MLflow
6d
Save
Mark Applied
Hide
Senior Engineering Manager - Containers at Edge
San Francisco or Denver or New York City
$228k-$274k/yr HybridFull Time
Fastly
FastlyNYSE: FSLY: Provides edge cloud platform for content delivery and cybersecurity.
10+ YOE4+ Mgmt4+ years managing engineering/SRE teams,10+ years technical experience with distributed systems/CDN/cloud infrastructure, hands-on development in Go/Rust/C/C++, Linux, roadmap and cross-functional leadership.
Go, Rust, C/C++, Linux
2w
Save
Mark Applied
Hide
Engineering Manager, Platform Infrastructure (Foundations)
San Francisco, California, United States
$270k-$320k/yr OnsiteFull Time
Anyscale
Anyscale: Cloud platform for scaling distributed machine learning applications.
Experienced engineering leader for infrastructure, SRE, and governance; deep distributed systems knowledge; Kubernetes and cloud provider experience; hiring, coaching, and execution skills.
Kubernetes, AWS, GCP, Azure, VMs
1mo
Save
Mark Applied
Hide
Engineering Manager, Reliability Platform
San Francisco or Sunnyvale or New York City
$194k-$285k/yr OnsiteFull Time
DoorDash
DoorDashNASDAQ: DASH: On-demand delivery platform connecting consumers with local merchants.
5+ YOE5+ Mgmt5+ years leading engineering teams and 5+ years in infrastructure/platform/backend roles; strong platform mindset, SRE experience (SLOs/error budgets), AWS and cloud fundamentals, influence and hiring experience, experience with incident/incident response processes.
AWS, Kafka, MCP, Covey Scout, Covey, IDE
1mo
Save
Mark Applied
Hide
Manager, Software Engineering (Reliability Platform)
United States or California or Washington or New York or New Jersey or Connecticut or Los Angeles or San Francisco
$204k-$290k/yr RemoteFull Time
Affirm
AffirmNasdaq: AFRM: Financial platform providing installment loans for consumer purchases.
7+ YOE2+ Mgmt7+ years backend/full-stack engineering experience with 2+ years engineering leadership; SRE/production engineering experience; observability and platform-building experience; strong programming (Python, Kotlin, Java); Bachelor\u0002s degree or equivalent experience.
Python, Kotlin, Java
3mo
Save
Mark Applied
Hide
Senior Engineering Manager, Service Enablement
San Francisco, California, United States
$212k-$318k/yr HybridFull Time
Hinge Health
Hinge HealthNYSE: HNGE: Digital provider of musculoskeletal care and physical therapy
10+ YOE4+ Mgmt10+ years in technology; 4+ years leading engineering teams; strong cloud infra (AWS, Kubernetes/EKS) and IaC (Terraform); proven platform reliability, SRE focus, and cost optimization; cross-functional leadership across geographies.
AWS, Kubernetes, EKS, Terraform, NestJS, TypeScript, GitHub Actions, NX, Okteto, Datadog, CI/CD, Helm, Infisical, Vault
1mo
Save
Mark Applied
Hide
VP – IT Operations
Emeryville, California, United States
$220k-$250k/yr HybridFull Time
Grocery Outlet
Grocery OutletNASDAQ: GO: Discount grocery retailer selling overstocked and closeout name-brand products.
10+ YOE5+ Mgmt10+ years in IT infrastructure/operations, 5+ years senior leadership, experience with hybrid cloud (AWS/Azure/GCP), SRE, ITSM/ITIL, vendor management, budgeting and retail technology (POS) preferred.
AWS, Azure, GCP, POS, SRE, ITSM, ITIL
2mo
Save
Mark Applied
Hide
Director of Site Reliability Engineering
San Francisco, California, United States
$210k-$310k/yr HybridFull Time
Stellar Development Foundation
Stellar Development Foundation: Developing and maintaining the open-source Stellar blockchain network.
3+ YOE3+ MgmtLead SRE team; define OKRs; coordinate with dev teams; coach; 3+ years SRE and 3+ years managing; Kubernetes; IaC; strong communication.
Terraform, Ansible, Puppet, Kubernetes
4d
Save
Mark Applied
Hide
Lead Site Reliability Engineer
Charlotte or Chandler or San Francisco or Columbus
$119k-$224k/yr HybridFull Time
Wells Fargo
Wells FargoNYSE: WFC: Global provider of banking, investment, and mortgage financial services.
5+ YOERequires 5+ years in systems engineering or architecture, 5+ years SRE, cloud observability experience, hosting platforms, and familiarity with DevOps, Agile, and IT service management.
Elasticsearch, Kibana, Kafka, Airflow, Logstash, Grafana, Elastic APM, Jaeger, Zipkin, AWS, OCP, Kubernetes, PKS, Azure, VMware, Unix, Linux, Windows, Jenkins, Maven, Gradle, Groovy, Artifactory, GIT, Harness IO, Spinnaker, Terraform, UDeploy, AIOPS, ServiceNow, Remedy, Big Panda, Netcool
1w
Save
Mark Applied
Hide
Director, Site Reliability Engineering
New York City or San Francisco or Dallas
$197k-$314k/yr HybridFull Time
Salesforce
SalesforceNYSE: CRM: Sells cloud-based customer relationship management and business software solutions.
10+ YOE5+ MgmtBachelor's in a technical field,10+ years engineering experience with 5+ years leading SRE/Platform teams; experience with observability, incident management, distributed systems, and cloud architecture.
AWS, New Relic, Splunk, Datadog, Sentry, Honeycomb, Grafana, Prometheus, OpenTelemetry

Explore Jobs

Expand Your Job Search