85 site reliability manager jobs at 58 companies in Berkeley, CA

PromotedHiringCafe
Founding Backend / Infra Engineer
Cupertino, CA, US
$160k-$300k/yr On-SiteFull Time
HiringCafe
HiringCafe: Building a 100× better job search engine to take on Indeed and LinkedIn.
Own the crawlers, pipelines, and infrastructure powering a real-time job search engine. Strong Node.js and Python fundamentals; bonus points for security and reverse-engineering chops.
Node.js, Python, Elasticsearch, Redis
2mo
Save
Mark Applied
Hide
Site Reliability Engineer
San Francisco or South San Francisco
$150k/yr OnsiteFull Time
VantageScore
VantageScore: Provides credit scoring and data analytics solutions.
5+ YOEExperienced Site Reliability Engineer with a DevSecOps focus; patch management, vulnerability remediation; AWS and CI/CD, security tooling.
AWS, EC2, ECS, Lambda, EKS, S3, RDS, IAM, VPC, CloudTrail, Config, GuardDuty, GitHub Actions, CodePipeline, Terraform, CloudFormation, AWS CDK, Kubernetes, Snyk, Wiz, Prisma Cloud, Kong, HashiCorp Vault, Secrets Manager, CloudWatch, Datadog, Grafana
1mo
Save
Mark Applied
Hide
Site Reliability Engineering Manager, Storage - Apple Services Engineering
Cupertino, California, United States
OnsiteFull Time
Apple
AppleNASDAQ: AAPL: Designs and sells consumer electronics, software, and online services.
Engineering manager for distributed storage systems and large-scale storage infrastructure; experience in distributed systems and site reliability engineering.
2w
Save
Mark Applied
Hide
Manager, Site Reliability Engineering
San Francisco, California, United States
$204k-$306k/yr HybridFull Time
Okta
OktaNASDAQ: OKTA: Provide secure identity management and authentication for enterprises.
3+ Mgmt3+ years technical leadership experience; experience with cloud-native architectures, Kubernetes, Terraform, CI/CD, observability platforms; strong software development and automation background; US Person status required.
Amazon Web Services (AWS), Kubernetes, Terraform, Grafana, Splunk, APM, CI/CD
3mo
Save
Mark Applied
Hide
Senior Manager, Site Reliability Engineering
Santa Clara, California, United States
$200k-$322k/yr OnsiteFull Time
NVIDIA
NVIDIANASDAQ: NVDA: Designs GPU-accelerated computing and artificial intelligence hardware.
12+ YOE5+ MgmtLead and manage global IT operations and SRE teams; implement AI-driven reliability and automation; strong leadership and communication.
2mo
Save
Mark Applied
Hide
Senior Site Reliability Engineer
San Francisco, California, United States
$167k-$226k/yr HybridFull Time
Drata
Drata: Automated security and compliance platform for businesses.
6+ YOE6+ years of Site Reliability Engineering or related cloud engineering experience; strong cloud, Terraform, Docker, Linux; Datadog monitoring; CI/CD with GitHub Actions; incident management; AI experience preferred.
Terraform, Docker, Git, Linux, Datadog, GitHub Actions, Python, Bash, AWS, ECS, Kubernetes, MySQL
3mo
Save
Mark Applied
Hide
Senior Manager, Site Reliability Engineering
Santa Clara, California, United States
$200k-$322k/yr OnsiteFull Time
NVIDIA
NVIDIANASDAQ: NVDA: Designs graphics processing units and artificial intelligence hardware.
12+ YOE5+ MgmtDegree in CS/ECE/Physics/Math/Engineering or equivalent, 12+ years SRE/ITSM experience, 5+ years managing global IT/service teams, expertise in incident/problem/change management, observability, AI/automation, and executive communication.
3mo
Save
Mark Applied
Hide
Senior Site Reliability Engineer
San Francisco, California, United States
$195k-$240k/yr HybridFull Time
You.com
You.com: AI-powered search engine and enterprise productivity platform.
2+ YOE2+ years in SRE, 3+ years AWS with EKS and CI/CD, Python/Bash, incident management, Prometheus/Grafana, and building reliable infra.
OpenTelemetry, Prometheus, Grafana, Python, Bash, Terraform, Git, AWS, EKS, CI/CD
6d
Save
Mark Applied
Hide
Site Reliability Engineer
San Francisco, California, United States
HybridFull Time
Runloop
Runloop: Provides infrastructure and secure sandboxes for AI agents.
5+ YOE5+ years software engineering experience with 3+ years in SRE/DevOps, strong Python or Go skills, containerization, cloud infra, monitoring, networking, Linux administration, on‑call and incident management.
AWS, GCP, Azure, Grafana, Prometheus, Datadog, Python, Go, Docker, Kubernetes, Terraform, Pulumi, Sentry, RUM, CI/CD
1mo
Save
Mark Applied
Hide
Site Reliability Engineer
United States or San Francisco or New York City
$101k-$199k/yr OnsiteFull Time
Microsoft
MicrosoftNASDAQ: MSFT: Develops software, services, devices, and cloud computing solutions.
1+ YOEMaster's or Bachelor's in CS/IT (or equivalent experience), 1+ years managing physical infrastructure, on-call experience, experience with large-scale cloud/distributed systems preferred, and ability to pass Microsoft security screening.
Azure, InfiniBand, GPUs
3d
Save
Mark Applied
Hide
Senior Site Reliability Engineer
Sunnyvale, California, United States
$90k-$180k/yr OnsiteFull Time
Abbott
AbbottNYSE: ABT: Manufactures medical devices, diagnostics, and nutritional health products.
Ensure reliability, scalability, and performance of a medical-device remote monitoring platform; expertise in cloud (Azure), Kubernetes, observability, automation, and incident management; bachelor's in a technical discipline.
Python, Go, Bash, PowerShell, Microsoft Azure, Azure Kubernetes Service (AKS), Azure Monitor, Azure DevOps, Azure Policy, Kubernetes, Docker, Prometheus, Grafana, ELK/EFK, Datadog, Linux
2w
Save
Mark Applied
Hide
Site Reliability Engineer
Santa Clara or St. Louis or Bangalore or London or Paris or Melbourne or Taipei or Tokyo
OnsiteFull Time
Netskope
NetskopeNASDAQ: NTSK: Cloud-native cybersecurity and data protection platform for enterprises.
3+ YOEBachelor's in CS/Engineering or equivalent; 3+ years building/managing complex systems (including 1-2 years SRE); experience with cloud services, microservices, availability/performance optimization, debugging, and strong communication.
Python, C, C++, Go, Rust, Docker, Kubernetes, AWS, GCP, KVM, OpenNebula, OpenStack, TCP/IP
1mo
Save
Mark Applied
Hide
Senior Site Reliability Engineer
Oakland, California, United States
$175k-$210k/yr HybridFull Time
Fivetran
Fivetran: Automates data movement into cloud data warehouses.
5+ YOE5+ years SaaS experience; managed Kubernetes, cloud platforms (AWS/GCP/Azure), Terraform/Ansible/ArgoCD; Python/Shell scripting, Linux admin, PostgreSQL; incident response and reliability engineering experience.
Kubernetes, EKS, AKS, GKE, PostgreSQL, ArgoCD, Terraform, Ansible, Python, Shell, Go, Java, AWS, GCP, Azure, Grafana, Buildkite, Temporal, Pulumi, Linux, VPN, PrivateLink, Private Service Connect (GCP)
5d
Save
Mark Applied
Hide
Site Reliability Engineer
Santa Clara, California, United States
$230k-$250k/yr OnsiteFull Time
Forward Networks
Forward Networks: Provides a digital twin platform for enterprise network management.
6+ YOE6+ years SRE/DevOps experience in SaaS/cloud, strong networking fundamentals, Kubernetes, observability (Prometheus/Grafana/Datadog/Splunk), Python/Bash automation, cloud and IaC (AWS/GCP/Azure, Terraform/Ansible), and incident response ownership.
Kubernetes, Prometheus, Grafana, Datadog, Splunk, Python, Bash, AWS, GCP, Azure, Terraform, Ansible
3d
Save
Mark Applied
Hide
Staff Site Reliability Engineer
Santa Clara, California, United States
$163k-$214k/yr OnsiteFull Time
IonQ
IonQNYSE: IONQ: Develops and sells trapped-ion quantum computers and cloud services.
7+ YOE7+ years production engineering experience; hands-on AWS/GCP reliability, observability and SLO ownership, incident command, resilience testing, and multi-team technical leadership.
AWS, GCP, Amazon Bedrock Agent Core
1mo
Save
Mark Applied
Hide
Sr. Site Reliability Engineer
Palo Alto or Palo Alto or Washington
$165k-$230k/yr OnsiteFull Time
SpaceX
SpaceX: Designs and launches advanced rockets and satellite internet constellations.
5+ YOE5+ years experience with Kubernetes and Linux, proficiency in Bash/Python, experience with infrastructure automation and large-scale server management; Top Secret/SCI clearance required or obtainable.
Kubernetes, Linux, Bash, Python, Bazel, Makefiles, Terraform, Ansible, TCP/IP
3mo
Save
Mark Applied
Hide
Site Reliability Engineer (SRE)
Madrid or Lisbon or San Francisco
€55k-€68k/yr RemoteFull Time
Air Apps
Air Apps: Develops AI-powered productivity and utility mobile applications.
4+ YOE4+ years in SRE/DevOps/System Eng; cloud platforms (AWS/Azure/GCP); observability tools; IaC; containers; Linux; incident management; scripting; security; on-call.
Prometheus, Grafana, Datadog, ELK, Terraform, CloudFormation, Pulumi, Docker, Kubernetes, Helm, Python, Go, Bash
1w
Save
Mark Applied
Hide
Principal Site Reliability Engineer
San Francisco or Toronto
OnsiteFull Time
Cerebras Systems
Cerebras SystemsNasdaq: CBRS: Manufactures specialized computer chips designed for AI.
15+ YOE15+ years in SRE/infrastructure/platform engineering with large-scale fleets; experience in capacity management, orchestration, observability, SLOs/SLIs, incident response, and cross-team architecture.
Wafer-Scale Engine (WSE), Bazel
3d
Save
Mark Applied
Hide
Senior Site Reliability Engineer
Sunnyvale or Sylmar
$90k-$180k/yr OnsiteFull Time
Abbott
AbbottNYSE: ABT: Provides medical devices, diagnostics, and science-based nutritional products.
Senior SRE with strong distributed systems, cloud (Azure), Kubernetes, observability, automation, incident management, and cross-functional communication skills for a medical device remote monitoring platform.
Python, Go, Bash, PowerShell, Microsoft Azure, Azure Kubernetes Service (AKS), Azure Monitor, Azure DevOps, Azure Policy, Kubernetes, Docker, Prometheus, Grafana, ELK, EFK, Datadog, Linux
1mo
Save
Mark Applied
Hide
Staff Site Reliability Engineer
Foster City, California, United States
$250k-$300k/yr HybridFull Time
Zoox
ZooxNASDAQ: AMZN: Developing autonomous robotaxis for urban ride-hailing services.
5+ YOE5+ years operating GitHub Enterprise at scale, monorepo management, CI/CD integration, infrastructure-as-code (Terraform/Pulumi), cloud platform experience, technical leadership and migration planning.
Git, GitHub Enterprise, GitHub Cloud, Buildkite, GitHub Actions, Jenkins, GitLab CI, Terraform, Pulumi, Bazel, Buck, Reviewable, Gerrit
1mo
Save
Mark Applied
Hide
Site Reliability/Devops Engineer
San Francisco, California, United States
$100k-$200k/yr OnsiteFull Time
Graphon
Graphon: Developing graph-native AI models for multimodal data reasoning.
Proficient in Bash and Python; experience with infrastructure-as-code, Docker, CI/CD, multi-cloud deployments, networking and identity access; comfortable managing production environments and using AI tools.
Bash, Python, Infrastructure-as-code, Docker, CI/CD, AI tools