77 sre manager jobs at 55 companies in Fairfax, CA

1w
Save
Mark Applied
Hide
Senior Director – Observability | SRE
San Francisco or Dallas
OnsiteFull Time
Gap Inc.
Gap Inc.NYSE: GAP: A global specialty apparel with brands including Gap, Old Navy, Banana Republic, and Athleta.
10+ YOEStrategic technology leader with 10+ years driving operational transformation, expertise in ITIL, SRE, architecture, observability, infrastructure, cloud operations, automation, and service management, plus large-team leadership experience.
ITIL, SRE, Live Sight Insights
2d
Save
Mark Applied
Hide
SRE - Enterprise & Cloud Security - AI Driven Security - Senior Manager
New York City or Atlanta or Chicago or Washington or Boston or Dallas or San Francisco or Seattle or Houston
$124k-$280k/yr OnsiteFull Time
PwC
PwC: Global professional services network providing audit, tax, and consulting services.
7+ YOEBachelor's degree and 7+ years of experience required. Preferred fields include data science, AI, computer science, information systems, or engineering; relevant cloud, data engineering, or machine learning certifications preferred.
Python, Java, C++, TensorFlow, Scikit-Learn, AWS, Google Cloud, Microsoft Azure, Databricks, Snowflake
1w
Save
Mark Applied
Hide
Head of Engineering, Infrastructure & SRE
San Francisco or Boston or New York City or Austin or Tokyo or London or Bangalore or United States or India or Europe
OnsiteFull Time
Postman
Postman: Private software providing an API development and lifecycle-management platform for developers and organizations.
15+ YOE7+ MgmtRequires 15+ years in infrastructure, platform, or SRE engineering and 7+ years in leadership, with distributed-team management, Kubernetes ecosystem, Istio, AWS, Azure, and SRE expertise.
Kubernetes, Cluster API, Argo, Helm, Crossplane, Istio, AWS, Azure, GitOps, CI/CD, FinOps
4d
Save
Mark Applied
Hide
Sr Manager - Infrastructure, SRE, & AI Platforms - Services Special Projects
Cupertino or San Francisco
OnsiteFull Time
Apple
AppleNASDAQ: AAPL: Designing and manufacturing consumer electronics, software, and digital services.
12+ YOE6+ MgmtMS in computer science or related field; 12+ years engineering leadership; 6+ years managing multi-layered engineering organizations; expertise in cloud infrastructure, Kubernetes, AI/ML systems, SRE, and global operations.
Kubernetes, AWS, GCP, Cassandra, FoundationDB, Kafka, Redis, PostgreSQL
3w
Save
Mark Applied
Hide
Principal Product Manager Lead
United States or Seattle or San Francisco or Sunnyvale or Raleigh or Boston or London or Lisbon or Bengaluru or Dublin or Kyiv or Chicago or New York City or Austin or Phoenix or Las Vegas or Washington or Dallas
$200k-$275k/yr OnsiteFull Time
SingleStore
SingleStore: SingleStore is a private database software serving enterprises with real-time SQL for AI, applications, and analytics.
Product leadership experience with distributed databases, query processing, AI agents or MCP servers, observability systems, SRE workflows, complex technical initiatives, and cross-functional people leadership.
Aura Analyst, AI Agents, MCP servers, Python UDFs, Cloud Functions, Container Services, SQrL, SRE, CPU, GPU
2mo
Save
Mark Applied
Hide
Senior SRE Engineer - San Francisco
San Francisco, California, United States
HybridFull Time
Plaud
Plaud: AI note-taking hardware and software serving professionals with voice recording, transcription, and meeting-summary tools.
8+ YOE8+ years in SRE/Infrastructure/Platform engineering, strong cloud (AWS/GCP/Azure) and Kubernetes experience, on-call/incident management experience, proficiency in Go/Python/Java, and experience building observability and SLO-driven systems.
AWS, GCP, Azure, Kubernetes, Go, Python, Java, Cursor, GPT models, Gemini, Claude
2mo
Save
Mark Applied
Hide
Head - SRE
Bengaluru or Las Vegas or San Francisco
HybridFull Time
Skillz
SkillzNYSE: FIRY: Mobile esports platform for competitive gaming and player experiences.
14+ YOE14+ years infrastructure engineering experience with public cloud (AWS), 5+ years running Kubernetes (EKS), leadership experience, observability/CI-CD expertise, cost optimization track record, and proficiency in Go, Python, or Java.
EKS, EC2, VPC, IAM, Cost Explorer, Savings Plans, Kubernetes, Istio, Datadog, Prometheus, Jaeger, X-Ray, ArgoCD, GitHub Actions, Go, Python, Java
3mo
Save
Mark Applied
Hide
Senior Product Manager
Toronto or San Francisco
HybridFull Time
PagerDuty
PagerDutyNYSE: PD: Public software providing AI-powered digital operations and incident management software to business teams.
5+ YOE5+ years product management; SaaS B2B; SRE/DevOps domain experience; workflow automation and integration tooling; security and permissions expertise; strong communication; leadership in roadmap execution.
Workflow automation, APIs, Telemetry analysis, Security and authorization models, SaaS platforms
6h
Save
Mark Applied
Hide
Staff Product Manager, AI Infrastructure (Storage)
San Francisco or Sunnyvale
$200k-$240k/yr OnsiteFull Time
Crusoe
Crusoe: Vertically integrated AI infrastructure and energy.
5+ YOEBachelor's degree in electrical engineering, computer science, or related field; 5+ years of technical product management or product-minded engineering experience; cloud and storage product expertise required.
IaaS, PaaS, SaaS, SRE
3w
Save
Mark Applied
Hide
Principal Product Manager Lead
United States or Seattle or San Francisco or Sunnyvale or Raleigh or Boston or London or Lisbon or Bangalore or Dublin or Kyiv or Chicago or New York City or Austin or Phoenix or Las Vegas or Dallas
$200k-$275k/yr OnsiteFull Time
SingleStore
SingleStore: Private database software serving enterprises with real-time SQL for AI, applications, and analytics.
Requires deep distributed database, query processing, performance optimization, observability, and AI product expertise, plus multi-team leadership, customer engagement, execution, communication, and mentorship skills.
Aura Analyst, AI Agents, MCP servers, Python UDFs, Cloud Functions, Container Services, SQrL, SRE
3d
Save
Mark Applied
Hide
Director, Technical Program Manager (Resiliency and Reliability Engineering)
McLean or Richmond or New York City or Plano or San Francisco
$210k-$287k/yr OnsiteFull Time
Capital One
Capital OneNYSE: COF: A technology-driven bank providing diverse financial services.
7+ YOEBachelor's degree and 7+ years managing technical programs required. Preferred: distributed systems, cloud, SRE, resilience engineering, Agile delivery, complex program leadership, and regulated-environment experience.
Agile
1w
Save
Mark Applied
Hide
Administrative Business Partner, YouTube SRE
San Bruno, California, United States
$96k-$137k/yr OnsiteFull Time
YouTube
YouTube: Google-owned video-sharing and content-distribution platform serving viewers, creators, and advertisers.
3+ YOERequires 3 years of administrative experience in a technology or multinational environment. Preferred: 6 years supporting executives and managing small-scale projects and events.
Google products and services
3w
Save
Mark Applied
Hide
Senior Staff TPM - Core Infrastructure & Platform Evolution
Palo Alto or Virginia or Miami or São Paulo
HybridFull Time
Nubank
NubankNYSE: NU: Brazilian digital bank serving consumers and small businesses with credit cards, accounts, payments, lending, investments, and insurance.
Requires deep infrastructure, SRE, or backend engineering expertise; hands-on technical leadership; massive cloud-native systems experience; architectural mastery; AI/ML transformation experience; and executive communication skills.
AI, ML, SRE, cloud-native, distributed databases, microservices
2w
Save
Mark Applied
Hide
Technical Customer Success / Customer Program Manager
United States or Pleasanton or India
RemoteFull Time
Ciroos
Ciroos: Private AI SRE software helping enterprise SRE, DevOps, and operations teams investigate incidents and reduce toil.
Experience leading complex enterprise customer programs; technical fluency across cloud, SRE, observability, integrations, and security; excellent communication and customer relationship skills.
AWS, Microsoft Azure, Google Cloud Platform, Kubernetes, Splunk, Datadog, ServiceNow, Slack, Jira, Linear, SRE, ITSM
3mo
Save
Mark Applied
Hide
Senior Technical Program Manager, Infrastructure
Mountain View, California, United States
$198k-$236k/yr HybridFull Time
Glean
Glean: Private enterprise AI software serving organizations with search, answers, content generation, and workflow automation.
8+ YOE8+ years TPM/infrastructure or SRE experience with 3+ years leading infra/platform programs; BS/MS in CS/Engineering or related; strong cloud, ML/LLM, reliability, and cross-functional leadership skills.
Microsoft Teams, Zoom, ServiceNow, Zendesk, GitHub, AWS, GCP, Azure, LLM
1w
Save
Mark Applied
Hide
Principal Technical Program Manager (TPM) - AI Infrastructure Operations
Houston or New York City or San Francisco or Seattle
OnsiteFull Time
Nscale
Nscale: Vertically integrated AI infrastructure and GPU cloud platform.
5+ YOE5+ years in technical program management for complex infrastructure or software programs; knowledge of data centers, distributed systems, Linux, networking, operational metrics, and Agile/Scrum. Technical degree preferred.
Linux, Agile, Scrum, NVIDIA GPUs, InfiniBand, RDMA, SRE, Continuous Integration/Continuous Deployment (CI/CD)
3w
Save
Mark Applied
Hide
IT Systems Engineer - Internal Platforms & SRE
San Francisco or San Jose
$206k-$275k/yr HybridFull Time
Lambda
Lambda: AI infrastructure building GPU cloud services and supercomputers for researchers, enterprises, and hyperscalers.
Experience with system design, scalable cloud infrastructure, configuration management, programming in Python or Go, distributed systems, automation, documentation, and cross-functional collaboration.
AWS, GCP, Azure, Chef, Ansible, Terraform, GitHub Actions, Python, Go
2mo
Save
Mark Applied
Hide
Engineering Manager, Cloud Platform
San Francisco, California, United States
$165k-$330k/yr HybridFull Time
Baseten
Baseten: Private AI infrastructure provider that trains, deploys, and serves artificial-intelligence models for businesses.
Experienced manager for cloud platform/SRE teams with strong Kubernetes and production infrastructure background, familiar with IaC and CI/CD tooling, recruiting and incident management skills.
Kubernetes, Terraform, CloudFormation, Pulumi, GitHub Actions, GitLab CI, CircleCI, Jenkins, Prometheus, ELK stack, Grafana, OpenTelemetry
1mo
Save
Mark Applied
Hide
Incident Manager
United States or Texas or San Francisco
$104k-$146k/yr RemoteFull Time
Databricks
Databricks: Data and AI software providing a unified platform.
5+ YOE5+ years in incident management/SRE/production operations for cloud-native systems; lead high-severity incidents; cloud (AWS/Azure/GCP) and observability expertise; log analysis; scripting (Python/Go/Bash); BS/MS in CS/CE or related field.
AWS, Azure, GCP, Datadog, Elasticsearch, Splunk, Cloud Logging, OpenTelemetry, Prometheus, Grafana, Python, Go, Bash, Apache Spark, Delta Lake, MLflow
2w
Save
Mark Applied
Hide
Engineering Manager, Kubernetes Infrastructure (Bare Metal)
Livingston or New York City or Sunnyvale or Bellevue or San Francisco
$182k-$242k/yr OnsiteFull Time
CoreWeave
CoreWeaveNasdaq: CRWV: Specialized cloud provider for large-scale AI and machine learning.
Experience managing infrastructure, platform, or SRE engineering teams; strong Kubernetes, distributed systems, production infrastructure, incident response, communication, and cross-functional leadership skills.
Kubernetes, Go, Python

Explore Jobs

Expand Your Job Search