160 site reliability manager jobs at 53 companies in Tiburon, CA

3mo
Save
Mark Applied
Hide
Site Reliability Engineer
San Francisco or South San Francisco
$150k/yr OnsiteFull Time
VantageScore Solutions, LLC
VantageScore Solutions, LLC: A Higher Level of Confidence
5+ YOEExperienced Site Reliability Engineer with a DevSecOps focus; patch management, vulnerability remediation; AWS and CI/CD, security tooling.
AWS, EC2, ECS, Lambda, EKS, S3, RDS, IAM, VPC, CloudTrail, Config, GuardDuty, GitHub Actions, CodePipeline, Terraform, CloudFormation, AWS CDK, Kubernetes, Snyk, Wiz, Prisma Cloud, Kong, HashiCorp Vault, Secrets Manager, CloudWatch, Datadog, Grafana
1mo
Save
Mark Applied
Hide
Manager, Site Reliability Engineering
San Francisco, California, United States
$204k-$306k/yr HybridFull Time
Okta
OktaNASDAQ: OKTA: Identity management and access control software provider.
3+ Mgmt3+ years technical leadership experience; experience with cloud-native architectures, Kubernetes, Terraform, CI/CD, observability platforms; strong software development and automation background; US Person status required.
Amazon Web Services (AWS), Kubernetes, Terraform, Grafana, Splunk, APM, CI/CD
1w
Save
Mark Applied
Hide
Senior Manager, Site Reliability Engineering
Santa Clara, California, United States
$248k-$397k/yr OnsiteFull Time
NVIDIA
NVIDIANASDAQ: NVDA: Computing platform for AI and accelerated graphics.
12+ YOE5+ MgmtBachelor's, master's, or doctoral degree in a related field or equivalent experience; 12+ years in SRE or IT service management and 5+ years leading global IT operations or service management teams.
AI, SRE, SLOs, ITIL
1w
Save
Mark Applied
Hide
Senior Manager, Site Reliability Engineering
Mountain View or Mountain View or California or United States
$222k-$301k/yr OnsiteFull Time
Intuit
IntuitNASDAQ: INTU: A global financial technology platform powering prosperity.
8+ YOE3+ Mgmt8+ years in systems, SRE, or infrastructure engineering; 3+ years managing engineering teams; AWS at scale; distributed systems, Kubernetes, IaC, observability, incident management, and AI Ops experience; bachelor's degree required.
AWS, Amazon EC2, Amazon EKS, Amazon ECS, Amazon VPC, Amazon RDS, Amazon DynamoDB, AWS IAM, Amazon CloudWatch, AWS Auto Scaling, Kubernetes, Terraform, AWS CloudFormation, Datadog, Splunk, PagerDuty, Prometheus, Grafana, AI Ops, AIOps platforms
1mo
Save
Mark Applied
Hide
Site Reliability Engineer
San Francisco, California, United States
HybridFull Time
Runloop AI
Runloop AI: Runloop AI provides AI infrastructure, secure code sandboxes, and evaluation tools for developers building software-engineering agents.
5+ YOE5+ years software engineering experience with 3+ years in SRE/DevOps, strong Python or Go skills, containerization, cloud infra, monitoring, networking, Linux administration, on‑call and incident management.
AWS, GCP, Azure, Grafana, Prometheus, Datadog, Python, Go, Docker, Kubernetes, Terraform, Pulumi, Sentry, RUM, CI/CD
1mo
Save
Mark Applied
Hide
Senior Site Reliability Engineer
Sunnyvale, California, United States
$90k-$180k/yr OnsiteFull Time
Abbott
AbbottNYSE: ABT: Manufactures medical devices, diagnostics, and nutritional health products.
Ensure reliability, scalability, and performance of a medical-device remote monitoring platform; expertise in cloud (Azure), Kubernetes, observability, automation, and incident management; bachelor's in a technical discipline.
Python, Go, Bash, PowerShell, Microsoft Azure, Azure Kubernetes Service (AKS), Azure Monitor, Azure DevOps, Azure Policy, Kubernetes, Docker, Prometheus, Grafana, ELK/EFK, Datadog, Linux
2mo
Save
Mark Applied
Hide
Senior Site Reliability Engineer
Oakland, California, United States
$175k-$210k/yr HybridFull Time
Fivetran
Fivetran: Automated data movement and integration platform for organizations.
5+ YOE5+ years SaaS experience; managed Kubernetes, cloud platforms (AWS/GCP/Azure), Terraform/Ansible/ArgoCD; Python/Shell scripting, Linux admin, PostgreSQL; incident response and reliability engineering experience.
Kubernetes, EKS, AKS, GKE, PostgreSQL, ArgoCD, Terraform, Ansible, Python, Shell, Go, Java, AWS, GCP, Azure, Grafana, Buildkite, Temporal, Pulumi, Linux, VPN, PrivateLink, Private Service Connect (GCP)
1mo
Save
Mark Applied
Hide
Site Reliability Engineer
Santa Clara, California, United States
$230k-$250k/yr OnsiteFull Time
Forward Networks
Forward Networks: Private network-software helping enterprises and government agencies model, secure, and automate complex hybrid networks.
6+ YOE6+ years SRE/DevOps experience in SaaS/cloud, strong networking fundamentals, Kubernetes, observability (Prometheus/Grafana/Datadog/Splunk), Python/Bash automation, cloud and IaC (AWS/GCP/Azure, Terraform/Ansible), and incident response ownership.
Kubernetes, Prometheus, Grafana, Datadog, Splunk, Python, Bash, AWS, GCP, Azure, Terraform, Ansible
3w
Save
Mark Applied
Hide
Senior Site Reliability Engineer - Cloud
Santa Clara, California, United States
$168k-$265k/yr OnsiteFull Time
NVIDIA
NVIDIANASDAQ: NVDA: Computing platform for AI and accelerated graphics.
8+ YOEMS/BS or equivalent experience, 8+ years supporting live-site production, SRE on-call experience, strong Kubernetes and Python skills, Akamai/CDN and AWS experience, incident management and automation focus.
Akamai Edge Redirector Cloudlets, Akamai Forward Rewrite Cloudlets, Akamai Cloudlets Policy Manager, Akamai CDN, WAF, AWS, Kubernetes, Python
2d
Save
Mark Applied
Hide
Site Reliability Engineering Manager, Vehicle Software
Sunnyvale, California, United States
$276k-$294k/yr HybridFull Time
Wayve
Wayve: British autonomous-driving software licensing vehicle-agnostic AI Driver technology to automakers and fleet owners.
8+ YOE3+ MgmtRequires 8+ years building production software systems, 3+ years people leadership, SRE and reliability practices, architecture experience, and production coding in C++, Rust, Python, or Go.
Linux, C++, Rust, Python, Go, CI/CD
2d
Save
Mark Applied
Hide
Site Reliability Engineering Manager, Vehicle Software
Sunnyvale, California, United States
$276k-$294k/yr HybridFull Time
Wayve
Wayve: British autonomous-driving software licensing vehicle-agnostic AI Driver technology to automakers and fleet owners.
8+ YOE3+ MgmtRequires 8+ years building production software systems, 3+ years of people leadership, SRE and reliability expertise, architecture experience, and hands-on coding in C++, Rust, Python, or Go.
C++, Rust, Python, Go, Linux, CI/CD
4w
Save
Mark Applied
Hide
Site Reliability Engineering (SRE) Manager, Apple Maps
Cupertino, California, United States
OnsiteFull Time
Apple
AppleNASDAQ: AAPL: Designing and manufacturing consumer electronics, software, and digital services.
Build, manage, and deliver highly available, automated infrastructure for Apple Maps at global scale; focus on reliability, scalability, and operational excellence.
1mo
Save
Mark Applied
Hide
Senior Site Reliability Engineer
Sunnyvale or Sylmar
$90k-$180k/yr OnsiteFull Time
Abbott
AbbottNYSE: ABT: Provides medical devices, diagnostics, and science-based nutritional products.
Senior SRE with strong distributed systems, cloud (Azure), Kubernetes, observability, automation, incident management, and cross-functional communication skills for a medical device remote monitoring platform.
Python, Go, Bash, PowerShell, Microsoft Azure, Azure Kubernetes Service (AKS), Azure Monitor, Azure DevOps, Azure Policy, Kubernetes, Docker, Prometheus, Grafana, ELK, EFK, Datadog, Linux
1mo
Save
Mark Applied
Hide
Principal Site Reliability Engineer
San Francisco or Toronto
OnsiteFull Time
Cerebras Systems
Cerebras SystemsNasdaq Global Select Market: CBRS: Designs processors and systems for AI training and inference.
15+ YOE15+ years in SRE/infrastructure/platform engineering with large-scale fleets; experience in capacity management, orchestration, observability, SLOs/SLIs, incident response, and cross-team architecture.
Wafer-Scale Engine (WSE), Bazel
2mo
Save
Mark Applied
Hide
Staff Site Reliability Engineer
Foster City, California, United States
$250k-$300k/yr HybridFull Time
Zoox
ZooxNASDAQ: AMZN: Developing autonomous robotaxis for urban ride-hailing services.
5+ YOE5+ years operating GitHub Enterprise at scale, monorepo management, CI/CD integration, infrastructure-as-code (Terraform/Pulumi), cloud platform experience, technical leadership and migration planning.
Git, GitHub Enterprise, GitHub Cloud, Buildkite, GitHub Actions, Jenkins, GitLab CI, Terraform, Pulumi, Bazel, Buck, Reviewable, Gerrit
2w
Save
Mark Applied
Hide
K8 Site Reliability SME
San Jose or Austin
RemoteFull Time
Bitdeer
BitdeerThe Nasdaq Stock Market LLC: BTDR: Public Singaporean Bitcoin mining and AI cloud infrastructure serving enterprises with computing, datacenters, and mining solutions.
5+ YOERequires 5+ years of Kubernetes operations, 2+ years managing GPU workloads, Terraform, Helm, GitOps, SRE practices, monitoring, Go or Python, and multi-tenant platform experience.
Kubernetes, Nvidia GPU operator, Terraform, Helm, ArgoCD, Flux, Prometheus, Grafana, Alertmanager, PagerDuty, Go, Python, Slurm, Ray, Kubeflow, Ironic, MAAS, GitOps
2mo
Save
Mark Applied
Hide
Site Reliability/Devops Engineer
San Francisco, California, United States
$100k-$200k/yr OnsiteFull Time
Graphon AI
Graphon AI: Graphon AI is a private enterprise-AI software building multimodal relational memory for organizations and AI agents.
Proficient in Bash and Python; experience with infrastructure-as-code, Docker, CI/CD, multi-cloud deployments, networking and identity access; comfortable managing production environments and using AI tools.
Bash, Python, Infrastructure-as-code, Docker, CI/CD, AI tools
1mo
Save
Mark Applied
Hide
Senior Site Reliability Engineer
San Francisco, California, United States
$149k-$224k/yr HybridFull Time
Salesforce
SalesforceNYSE: CRM: Sells cloud-based customer relationship management and business software solutions.
5+ YOE5+ years systems and software engineering experience for large-scale internet services; expertise in SRE principles, containers, observability, incident management, Python and Go, and applying AI/ML to operations.
Temporal, Airflow, Argo Workflows, Docker, Kubernetes, DNS, HTTP, Grafana, Prometheus, ELK, Splunk, Datadog, Python, Go, Linux, Claude Code, GitHub Copilot, Codex, Cursor, AWS, GCP, MCP
2mo
Save
Mark Applied
Hide
Sr. Site Reliability Engineer
Sunnyvale, California, United States
$170k-$196k/yr OnsiteFull Time
Illumio
Illumio: Private cybersecurity providing breach containment and Zero Trust security for enterprise organizations.
5+ YOE5+ years SRE experience with AWS and/or Azure, scripting in PowerShell/Python/Go, CI/CD experience (Azure DevOps, Jenkins, GitLab CI/CD), containerization knowledge (Docker, Kubernetes), bachelor’s degree or equivalent, on-call and incident management experience.
AWS, Azure, PowerShell, Python, Go, Azure DevOps, Jenkins, GitLab CI/CD, Docker, Kubernetes
3mo
Save
Mark Applied
Hide
Staff Site Reliability Engineer
Mountain View, California, United States
$252k-$308k/yr HybridFull Time
EarnIn
EarnIn: Fintech helping workers access earned wages in real time and manage finances without interest or mandatory fees.
7+ YOE7+ years in SRE or related field; experience applying AI/LLMs to operations; strong SLO/SLI and incident management; software engineering in Python or Go; observability and IaC proficiency; AI-assisted development tools; fintech/regulated environment experience.
Datadog, CloudWatch, OpenTelemetry, Terraform, Kubernetes, AWS, Python, Go, Cursor, Claude Code, Copilot

Explore Jobs

Expand Your Job Search