80 sre manager jobs at 61 companies in San Rafael, CA
2mo
Save
Mark Applied
Hide
2mo
Manager, Engineering - Dev Ops/SRE (Hybrid)
Sunnyvale, California, United States
$140k-$215k/yrHybridFull Time
CrowdStrikeNASDAQ: CRWD: Provides cloud-native endpoint protection and cybersecurity services.
10+ YOE3+ Mgmt10+ years software engineering experience with SRE/platform focus, 3+ years managing SRE teams, proficiency in cloud (AWS/Azure/GCP), deep SRE principles, incident management, bachelor's in CS or equivalent, able to work 2+ days/week in Sunnyvale.
United States or Seattle or San Francisco or Sunnyvale or Raleigh or Boston or London or Lisbon or Bengaluru or Dublin or Kyiv or Chicago or New York City or Austin or Phoenix or Las Vegas or Washington or Dallas
$200k-$275k/yrOnsiteFull Time
SingleStore: Provides a distributed SQL database for real-time analytics.
Product leadership experience with distributed databases, query processing, AI agents or MCP servers, observability systems, SRE workflows, complex technical initiatives, and cross-functional people leadership.
Plaud: Develops AI-powered voice recorders and automated transcription software.
8+ YOE8+ years in SRE/Infrastructure/Platform engineering, strong cloud (AWS/GCP/Azure) and Kubernetes experience, on-call/incident management experience, proficiency in Go/Python/Java, and experience building observability and SLO-driven systems.
SkillzNYSE: SKLZ: Operates a platform for competitive multiplayer mobile gaming.
14+ YOE14+ years infrastructure engineering experience with public cloud (AWS), 5+ years running Kubernetes (EKS), leadership experience, observability/CI-CD expertise, cost optimization track record, and proficiency in Go, Python, or Java.
Lambda: Provides high-performance GPU cloud infrastructure for AI development.
7+ YOE7+ years product management experience including 3+ years on technical infrastructure or platform products; experience defining observability, SLIs/SLOs/SLA, partnering with SRE and fleet engineering; strong writing, influence, and execution skills.
Sentry: Developer platform for error tracking and performance monitoring.
8+ YOE8+ years TPM experience in high-growth tech with platform/infrastructure/SRE/security exposure; proven cross-functional program leadership; incident management experience; strong analytical and communication skills.
WalmartNYSE: WMT: Multinational retail operating discount stores and supermarkets.
5+ YOEPlan and execute medium- to high-complexity technical programs, manage dependencies and risks, translate product needs into technical designs, apply SRE and cloud best practices, and align stakeholders across engineering and product teams.
United States or Seattle or San Francisco or Sunnyvale or Raleigh or Boston or London or Lisbon or Bangalore or Dublin or Kyiv or Chicago or New York City or Austin or Phoenix or Las Vegas or Dallas
$200k-$275k/yrOnsiteFull Time
SingleStore: Real-time distributed SQL database for transactions and analytics.
Requires deep distributed database, query processing, performance optimization, observability, and AI product expertise, plus multi-team leadership, customer engagement, execution, communication, and mentorship skills.
NubankNYSE: NU: Digital financial platform offering banking, credit, and investment services.
Requires deep infrastructure, SRE, or backend engineering expertise; hands-on technical leadership; massive cloud-native systems experience; architectural mastery; AI/ML transformation experience; and executive communication skills.
Technical Operational Program Manager - AI Infrastructure
Los Altos, California, United States
HybridFull Time
Majestic Labs: Developing memory-first AI server platforms for data centers.
10+ YOERequires 10+ years in program or project management, including 5+ years leading complex technical programs, manufacturing scale-up experience, infrastructure fluency, and strong communication skills.
Sr. Technical Program Manager, Product Escalations
Sunnyvale, California, United States
$157k-$180k/yrOnsiteFull Time
Illumio: Provides zero-trust segmentation software to contain cyberattacks.
7+ YOERequires 7+ years in technical support, escalation, account, incident, customer success engineering, or technical program management; enterprise escalation, SaaS/cloud, cross-functional, RCA, and executive communication experience.
Technical Customer Success / Customer Program Manager
United States or Pleasanton or India
RemoteFull Time
Ciroos: AI platform for automated site reliability engineering.
Experience leading complex enterprise customer programs; technical fluency across cloud, SRE, observability, integrations, and security; excellent communication and customer relationship skills.
AWS, Microsoft Azure, Google Cloud Platform, Kubernetes, Splunk, Datadog, ServiceNow, Slack, Jira, Linear, SRE, ITSM
Glean: AI platform for enterprise search and automated workplace agents
8+ YOE8+ years TPM/infrastructure or SRE experience with 3+ years leading infra/platform programs; BS/MS in CS/Engineering or related; strong cloud, ML/LLM, reliability, and cross-functional leadership skills.
Microsoft Teams, Zoom, ServiceNow, Zendesk, GitHub, AWS, GCP, Azure, LLM
Baseten: Scalable infrastructure platform for deploying and serving AI models.
Experienced manager for cloud platform/SRE teams with strong Kubernetes and production infrastructure background, familiar with IaC and CI/CD tooling, recruiting and incident management skills.
Wells FargoNYSE: WFC: Global provider of banking, investment, and mortgage financial services.
7+ YOE3+ MgmtRequires 7+ years in systems engineering or technology architecture, 3+ years of management, 5+ years leading engineering or SRE teams, and experience with customer-facing platforms, incident management, SRE, DevOps, cloud, and production operations.
Splunk, Grafana, AppDynamics, Dynatrace, OpenTelemetry, Prometheus, Kubernetes, OpenShift, AWS, Microsoft Azure, Google Cloud Platform, Infrastructure as Code (IaC), CI/CD, AIOps, ITIL
Bellevue or Livingston or New York City or Sunnyvale or San Francisco
$182k-$242k/yrOnsiteFull Time
CoreWeaveNASDAQ: CRWV: Cloud platform providing GPU-accelerated infrastructure for AI workloads.
7+ YOE3+ Mgmt3+ years engineering management and 7+ years technical experience in cloud operations/SRE; strong knowledge of cloud platforms, K8S, observability, incident management, and systems programming (Go); experience defining SLAs/SLOs and running on-call.
Databricks: A unified platform for data analytics and artificial intelligence.
5+ YOE5+ years in incident management/SRE/production operations for cloud-native systems; lead high-severity incidents; cloud (AWS/Azure/GCP) and observability expertise; log analysis; scripting (Python/Go/Bash); BS/MS in CS/CE or related field.
ThoughtSpot: AI-powered analytics platform for enterprise business intelligence.
5+ YOE5–8 years in TAM/SRE/technical support for enterprise SaaS; strong cloud (AWS/GCP/Azure), SQL, data warehousing, incident management, customer-facing communication and technical advisory skills.