77 sre manager jobs at 55 companies in Fairfax, CA
1w
Save
Mark Applied
Hide
1w
Senior Director – Observability | SRE
San Francisco or Dallas
OnsiteFull Time
Gap Inc.NYSE: GAP: A global specialty apparel with brands including Gap, Old Navy, Banana Republic, and Athleta.
10+ YOEStrategic technology leader with 10+ years driving operational transformation, expertise in ITIL, SRE, architecture, observability, infrastructure, cloud operations, automation, and service management, plus large-team leadership experience.
New York City or Atlanta or Chicago or Washington or Boston or Dallas or San Francisco or Seattle or Houston
$124k-$280k/yrOnsiteFull Time
PwC: Global professional services network providing audit, tax, and consulting services.
7+ YOEBachelor's degree and 7+ years of experience required. Preferred fields include data science, AI, computer science, information systems, or engineering; relevant cloud, data engineering, or machine learning certifications preferred.
Python, Java, C++, TensorFlow, Scikit-Learn, AWS, Google Cloud, Microsoft Azure, Databricks, Snowflake
San Francisco or Boston or New York City or Austin or Tokyo or London or Bangalore or United States or India or Europe
OnsiteFull Time
Postman: Private software providing an API development and lifecycle-management platform for developers and organizations.
15+ YOE7+ MgmtRequires 15+ years in infrastructure, platform, or SRE engineering and 7+ years in leadership, with distributed-team management, Kubernetes ecosystem, Istio, AWS, Azure, and SRE expertise.
Sr Manager - Infrastructure, SRE, & AI Platforms - Services Special Projects
Cupertino or San Francisco
OnsiteFull Time
AppleNASDAQ: AAPL: Designing and manufacturing consumer electronics, software, and digital services.
12+ YOE6+ MgmtMS in computer science or related field; 12+ years engineering leadership; 6+ years managing multi-layered engineering organizations; expertise in cloud infrastructure, Kubernetes, AI/ML systems, SRE, and global operations.
United States or Seattle or San Francisco or Sunnyvale or Raleigh or Boston or London or Lisbon or Bengaluru or Dublin or Kyiv or Chicago or New York City or Austin or Phoenix or Las Vegas or Washington or Dallas
$200k-$275k/yrOnsiteFull Time
SingleStore: SingleStore is a private database software serving enterprises with real-time SQL for AI, applications, and analytics.
Product leadership experience with distributed databases, query processing, AI agents or MCP servers, observability systems, SRE workflows, complex technical initiatives, and cross-functional people leadership.
Plaud: AI note-taking hardware and software serving professionals with voice recording, transcription, and meeting-summary tools.
8+ YOE8+ years in SRE/Infrastructure/Platform engineering, strong cloud (AWS/GCP/Azure) and Kubernetes experience, on-call/incident management experience, proficiency in Go/Python/Java, and experience building observability and SLO-driven systems.
SkillzNYSE: FIRY: Mobile esports platform for competitive gaming and player experiences.
14+ YOE14+ years infrastructure engineering experience with public cloud (AWS), 5+ years running Kubernetes (EKS), leadership experience, observability/CI-CD expertise, cost optimization track record, and proficiency in Go, Python, or Java.
Staff Product Manager, AI Infrastructure (Storage)
San Francisco or Sunnyvale
$200k-$240k/yrOnsiteFull Time
Crusoe: Vertically integrated AI infrastructure and energy.
5+ YOEBachelor's degree in electrical engineering, computer science, or related field; 5+ years of technical product management or product-minded engineering experience; cloud and storage product expertise required.
United States or Seattle or San Francisco or Sunnyvale or Raleigh or Boston or London or Lisbon or Bangalore or Dublin or Kyiv or Chicago or New York City or Austin or Phoenix or Las Vegas or Dallas
$200k-$275k/yrOnsiteFull Time
SingleStore: Private database software serving enterprises with real-time SQL for AI, applications, and analytics.
Requires deep distributed database, query processing, performance optimization, observability, and AI product expertise, plus multi-team leadership, customer engagement, execution, communication, and mentorship skills.
YouTube: Google-owned video-sharing and content-distribution platform serving viewers, creators, and advertisers.
3+ YOERequires 3 years of administrative experience in a technology or multinational environment. Preferred: 6 years supporting executives and managing small-scale projects and events.
NubankNYSE: NU: Brazilian digital bank serving consumers and small businesses with credit cards, accounts, payments, lending, investments, and insurance.
Requires deep infrastructure, SRE, or backend engineering expertise; hands-on technical leadership; massive cloud-native systems experience; architectural mastery; AI/ML transformation experience; and executive communication skills.
Technical Customer Success / Customer Program Manager
United States or Pleasanton or India
RemoteFull Time
Ciroos: Private AI SRE software helping enterprise SRE, DevOps, and operations teams investigate incidents and reduce toil.
Experience leading complex enterprise customer programs; technical fluency across cloud, SRE, observability, integrations, and security; excellent communication and customer relationship skills.
AWS, Microsoft Azure, Google Cloud Platform, Kubernetes, Splunk, Datadog, ServiceNow, Slack, Jira, Linear, SRE, ITSM
Glean: Private enterprise AI software serving organizations with search, answers, content generation, and workflow automation.
8+ YOE8+ years TPM/infrastructure or SRE experience with 3+ years leading infra/platform programs; BS/MS in CS/Engineering or related; strong cloud, ML/LLM, reliability, and cross-functional leadership skills.
Microsoft Teams, Zoom, ServiceNow, Zendesk, GitHub, AWS, GCP, Azure, LLM
Principal Technical Program Manager (TPM) - AI Infrastructure Operations
Houston or New York City or San Francisco or Seattle
OnsiteFull Time
Nscale: Vertically integrated AI infrastructure and GPU cloud platform.
5+ YOE5+ years in technical program management for complex infrastructure or software programs; knowledge of data centers, distributed systems, Linux, networking, operational metrics, and Agile/Scrum. Technical degree preferred.
Lambda: AI infrastructure building GPU cloud services and supercomputers for researchers, enterprises, and hyperscalers.
Experience with system design, scalable cloud infrastructure, configuration management, programming in Python or Go, distributed systems, automation, documentation, and cross-functional collaboration.
AWS, GCP, Azure, Chef, Ansible, Terraform, GitHub Actions, Python, Go
Baseten: Private AI infrastructure provider that trains, deploys, and serves artificial-intelligence models for businesses.
Experienced manager for cloud platform/SRE teams with strong Kubernetes and production infrastructure background, familiar with IaC and CI/CD tooling, recruiting and incident management skills.
Databricks: Data and AI software providing a unified platform.
5+ YOE5+ years in incident management/SRE/production operations for cloud-native systems; lead high-severity incidents; cloud (AWS/Azure/GCP) and observability expertise; log analysis; scripting (Python/Go/Bash); BS/MS in CS/CE or related field.