174 site reliability manager jobs at 56 companies in Danville, CA

3mo
Save
Mark Applied
Hide
Site Reliability Engineer
San Francisco or South San Francisco
$150k/yr OnsiteFull Time
VantageScore Solutions, LLC
VantageScore Solutions, LLC: A Higher Level of Confidence
5+ YOEExperienced Site Reliability Engineer with a DevSecOps focus; patch management, vulnerability remediation; AWS and CI/CD, security tooling.
AWS, EC2, ECS, Lambda, EKS, S3, RDS, IAM, VPC, CloudTrail, Config, GuardDuty, GitHub Actions, CodePipeline, Terraform, CloudFormation, AWS CDK, Kubernetes, Snyk, Wiz, Prisma Cloud, Kong, HashiCorp Vault, Secrets Manager, CloudWatch, Datadog, Grafana
3d
Save
Mark Applied
Hide
Manager, Site Reliability Engineer
San Francisco or New York City
$150k-$220k/yr OnsiteFull Time
Forge Global
Forge GlobalNYSE: FRGE: Financial technology operating a private-market marketplace and data, custody, and investment solutions for companies and investors.
10+ YOE5+ MgmtRequires 5+ years leading SRE, DevOps, cloud operations, or reliability functions; 10+ years in engineering or operations; bachelor's degree or equivalent; cloud infrastructure, distributed systems, observability, CI/CD, and automation experience.
AWS, Azure, Kubernetes, Terraform, Ansible, Datadog, CloudWatch
11h
Save
Mark Applied
Hide
Staff Site Reliability Engineer, Ads
San Francisco or United States
$217k-$304k/yr RemoteFull Time, Contract
Reddit
RedditNYSE: RDDT: Social news aggregation, web content rating, and discussion platform.
8+ YOE8+ years in site reliability or infrastructure engineering, distributed systems, cloud-native architecture, observability, automation, incident management, performance optimization, and backend software engineering.
Go, Kubernetes, Kafka, ClickHouse, Spark, Flink, BigQuery
1mo
Save
Mark Applied
Hide
Manager, Site Reliability Engineering
San Francisco, California, United States
$204k-$306k/yr HybridFull Time
Okta
OktaNASDAQ: OKTA: Identity management and access control software provider.
3+ Mgmt3+ years technical leadership experience; experience with cloud-native architectures, Kubernetes, Terraform, CI/CD, observability platforms; strong software development and automation background; US Person status required.
Amazon Web Services (AWS), Kubernetes, Terraform, Grafana, Splunk, APM, CI/CD
1w
Save
Mark Applied
Hide
Senior Manager, Site Reliability Engineering
Santa Clara, California, United States
$248k-$397k/yr OnsiteFull Time
NVIDIA
NVIDIANASDAQ: NVDA: Computing platform for AI and accelerated graphics.
12+ YOE5+ MgmtBachelor's, master's, or doctoral degree in a related field or equivalent experience; 12+ years in SRE or IT service management and 5+ years leading global IT operations or service management teams.
AI, SRE, SLOs, ITIL
1w
Save
Mark Applied
Hide
Senior Manager, Site Reliability Engineering
Mountain View or Mountain View or California or United States
$222k-$301k/yr OnsiteFull Time
Intuit
IntuitNASDAQ: INTU: A global financial technology platform powering prosperity.
8+ YOE3+ Mgmt8+ years in systems, SRE, or infrastructure engineering; 3+ years managing engineering teams; AWS at scale; distributed systems, Kubernetes, IaC, observability, incident management, and AI Ops experience; bachelor's degree required.
AWS, Amazon EC2, Amazon EKS, Amazon ECS, Amazon VPC, Amazon RDS, Amazon DynamoDB, AWS IAM, Amazon CloudWatch, AWS Auto Scaling, Kubernetes, Terraform, AWS CloudFormation, Datadog, Splunk, PagerDuty, Prometheus, Grafana, AI Ops, AIOps platforms
1mo
Save
Mark Applied
Hide
Site Reliability Engineer
San Francisco, California, United States
HybridFull Time
Runloop AI
Runloop AI: Runloop AI provides AI infrastructure, secure code sandboxes, and evaluation tools for developers building software-engineering agents.
5+ YOE5+ years software engineering experience with 3+ years in SRE/DevOps, strong Python or Go skills, containerization, cloud infra, monitoring, networking, Linux administration, on‑call and incident management.
AWS, GCP, Azure, Grafana, Prometheus, Datadog, Python, Go, Docker, Kubernetes, Terraform, Pulumi, Sentry, RUM, CI/CD
1mo
Save
Mark Applied
Hide
Senior Site Reliability Engineer
Sunnyvale, California, United States
$90k-$180k/yr OnsiteFull Time
Abbott
AbbottNYSE: ABT: Global healthcare technology focused on life-changing medical innovations.
Ensure reliability, scalability, and performance of a medical-device remote monitoring platform; expertise in cloud (Azure), Kubernetes, observability, automation, and incident management; bachelor's in a technical discipline.
Python, Go, Bash, PowerShell, Microsoft Azure, Azure Kubernetes Service (AKS), Azure Monitor, Azure DevOps, Azure Policy, Kubernetes, Docker, Prometheus, Grafana, ELK/EFK, Datadog, Linux
2mo
Save
Mark Applied
Hide
Senior Site Reliability Engineer
Oakland, California, United States
$175k-$210k/yr HybridFull Time
Fivetran
Fivetran: Automated data movement and integration platform for organizations.
5+ YOE5+ years SaaS experience; managed Kubernetes, cloud platforms (AWS/GCP/Azure), Terraform/Ansible/ArgoCD; Python/Shell scripting, Linux admin, PostgreSQL; incident response and reliability engineering experience.
Kubernetes, EKS, AKS, GKE, PostgreSQL, ArgoCD, Terraform, Ansible, Python, Shell, Go, Java, AWS, GCP, Azure, Grafana, Buildkite, Temporal, Pulumi, Linux, VPN, PrivateLink, Private Service Connect (GCP)
1mo
Save
Mark Applied
Hide
Site Reliability Engineer
Santa Clara, California, United States
$230k-$250k/yr OnsiteFull Time
Forward Networks
Forward Networks: Private network-software helping enterprises and government agencies model, secure, and automate complex hybrid networks.
6+ YOE6+ years SRE/DevOps experience in SaaS/cloud, strong networking fundamentals, Kubernetes, observability (Prometheus/Grafana/Datadog/Splunk), Python/Bash automation, cloud and IaC (AWS/GCP/Azure, Terraform/Ansible), and incident response ownership.
Kubernetes, Prometheus, Grafana, Datadog, Splunk, Python, Bash, AWS, GCP, Azure, Terraform, Ansible
3w
Save
Mark Applied
Hide
Senior Site Reliability Engineer - Cloud
Santa Clara, California, United States
$168k-$265k/yr OnsiteFull Time
NVIDIA
NVIDIANASDAQ: NVDA: Computing platform for AI and accelerated graphics.
8+ YOEMS/BS or equivalent experience, 8+ years supporting live-site production, SRE on-call experience, strong Kubernetes and Python skills, Akamai/CDN and AWS experience, incident management and automation focus.
Akamai Edge Redirector Cloudlets, Akamai Forward Rewrite Cloudlets, Akamai Cloudlets Policy Manager, Akamai CDN, WAF, AWS, Kubernetes, Python
3d
Save
Mark Applied
Hide
Site Reliability Engineering Manager, Vehicle Software
Sunnyvale, California, United States
$276k-$294k/yr HybridFull Time
Wayve
Wayve: British autonomous-driving software licensing vehicle-agnostic AI Driver technology to automakers and fleet owners.
8+ YOE3+ MgmtRequires 8+ years building production software systems, 3+ years people leadership, SRE and reliability practices, architecture experience, and production coding in C++, Rust, Python, or Go.
Linux, C++, Rust, Python, Go, CI/CD
3d
Save
Mark Applied
Hide
Site Reliability Engineering Manager, Vehicle Software
Sunnyvale, California, United States
$276k-$294k/yr HybridFull Time
Wayve
Wayve: British autonomous-driving software licensing vehicle-agnostic AI Driver technology to automakers and fleet owners.
8+ YOE3+ MgmtRequires 8+ years building production software systems, 3+ years of people leadership, SRE and reliability expertise, architecture experience, and hands-on coding in C++, Rust, Python, or Go.
C++, Rust, Python, Go, Linux, CI/CD
4w
Save
Mark Applied
Hide
Site Reliability Engineering (SRE) Manager, Apple Maps
Cupertino, California, United States
OnsiteFull Time
Apple
AppleNASDAQ: AAPL: Designing and manufacturing consumer electronics, software, and digital services.
Build, manage, and deliver highly available, automated infrastructure for Apple Maps at global scale; focus on reliability, scalability, and operational excellence.
1mo
Save
Mark Applied
Hide
Principal Site Reliability Engineer
San Francisco or Toronto
OnsiteFull Time
Cerebras Systems
Cerebras SystemsNasdaq Global Select Market: CBRS: Designs processors and systems for AI training and inference.
15+ YOE15+ years in SRE/infrastructure/platform engineering with large-scale fleets; experience in capacity management, orchestration, observability, SLOs/SLIs, incident response, and cross-team architecture.
Wafer-Scale Engine (WSE), Bazel
1mo
Save
Mark Applied
Hide
Senior Site Reliability Engineer
Sunnyvale or Sylmar
$90k-$180k/yr OnsiteFull Time
Abbott
AbbottNYSE: ABT: Global healthcare technology focused on life-changing medical innovations.
Senior SRE with strong distributed systems, cloud (Azure), Kubernetes, observability, automation, incident management, and cross-functional communication skills for a medical device remote monitoring platform.
Python, Go, Bash, PowerShell, Microsoft Azure, Azure Kubernetes Service (AKS), Azure Monitor, Azure DevOps, Azure Policy, Kubernetes, Docker, Prometheus, Grafana, ELK, EFK, Datadog, Linux
1d
Save
Mark Applied
Hide
Sr. Site Reliability Engineer
Scottsdale or San Francisco or Chicago or New York City or Phoenix or Chicago or San Francisco
$118k-$183k/yr HybridFull Time
Early Warning Services
Early Warning Services: U.S. bank-owned fintech and consumer reporting agency providing identity, fraud-risk, and real-time payment solutions to financial institutions.
5+ YOEBachelor's degree in business, computer science, or related field; 5+ years of technical experience; incident management, Linux, scripting, observability, Git, security protocols, and enterprise production experience.
Observability, Linux, Git, Java, Ruby, Python, JavaScript, Go, CI/CD, TCP/UDP/IP, AWS, Docker, Kubernetes, Swarm
2mo
Save
Mark Applied
Hide
Site Reliability/Devops Engineer
San Francisco, California, United States
$100k-$200k/yr OnsiteFull Time
Graphon AI
Graphon AI: Graphon AI is a private enterprise-AI software building multimodal relational memory for organizations and AI agents.
Proficient in Bash and Python; experience with infrastructure-as-code, Docker, CI/CD, multi-cloud deployments, networking and identity access; comfortable managing production environments and using AI tools.
Bash, Python, Infrastructure-as-code, Docker, CI/CD, AI tools
2w
Save
Mark Applied
Hide
K8 Site Reliability SME
San Jose or Austin
RemoteFull Time
Bitdeer
BitdeerThe Nasdaq Stock Market LLC: BTDR: Public Singaporean Bitcoin mining and AI cloud infrastructure serving enterprises with computing, datacenters, and mining solutions.
5+ YOERequires 5+ years of Kubernetes operations, 2+ years managing GPU workloads, Terraform, Helm, GitOps, SRE practices, monitoring, Go or Python, and multi-tenant platform experience.
Kubernetes, Nvidia GPU operator, Terraform, Helm, ArgoCD, Flux, Prometheus, Grafana, Alertmanager, PagerDuty, Go, Python, Slurm, Ray, Kubeflow, Ironic, MAAS, GitOps
2mo
Save
Mark Applied
Hide
Staff Site Reliability Engineer
Foster City, California, United States
$250k-$300k/yr HybridFull Time
Zoox
Zoox: Autonomous mobility developing a fully electric robotaxi fleet.
5+ YOE5+ years operating GitHub Enterprise at scale, monorepo management, CI/CD integration, infrastructure-as-code (Terraform/Pulumi), cloud platform experience, technical leadership and migration planning.
Git, GitHub Enterprise, GitHub Cloud, Buildkite, GitHub Actions, Jenkins, GitLab CI, Terraform, Pulumi, Bazel, Buck, Reviewable, Gerrit

Explore Jobs

Expand Your Job Search