116 site reliability manager jobs at 42 companies in Larkspur, CA

3mo
Save
Mark Applied
Hide
Site Reliability Engineer
San Francisco or South San Francisco
$150k/yr OnsiteFull Time
VantageScore
VantageScore: Provides credit scoring and data analytics solutions.
5+ YOEExperienced Site Reliability Engineer with a DevSecOps focus; patch management, vulnerability remediation; AWS and CI/CD, security tooling.
AWS, EC2, ECS, Lambda, EKS, S3, RDS, IAM, VPC, CloudTrail, Config, GuardDuty, GitHub Actions, CodePipeline, Terraform, CloudFormation, AWS CDK, Kubernetes, Snyk, Wiz, Prisma Cloud, Kong, HashiCorp Vault, Secrets Manager, CloudWatch, Datadog, Grafana
1mo
Save
Mark Applied
Hide
Manager, Site Reliability Engineering
San Francisco, California, United States
$204k-$306k/yr HybridFull Time
Okta
OktaNASDAQ: OKTA: Provide secure identity management and authentication for enterprises.
3+ Mgmt3+ years technical leadership experience; experience with cloud-native architectures, Kubernetes, Terraform, CI/CD, observability platforms; strong software development and automation background; US Person status required.
Amazon Web Services (AWS), Kubernetes, Terraform, Grafana, Splunk, APM, CI/CD
1w
Save
Mark Applied
Hide
Senior Manager, Site Reliability Engineering
Mountain View or Mountain View or California or United States
$222k-$301k/yr OnsiteFull Time
Intuit
IntuitNASDAQ: INTU: Provides financial software for accounting, tax, and personal finance.
8+ YOE3+ Mgmt8+ years in systems, SRE, or infrastructure engineering; 3+ years managing engineering teams; AWS at scale; distributed systems, Kubernetes, IaC, observability, incident management, and AI Ops experience; bachelor's degree required.
AWS, Amazon EC2, Amazon EKS, Amazon ECS, Amazon VPC, Amazon RDS, Amazon DynamoDB, AWS IAM, Amazon CloudWatch, AWS Auto Scaling, Kubernetes, Terraform, AWS CloudFormation, Datadog, Splunk, PagerDuty, Prometheus, Grafana, AI Ops, AIOps platforms
1mo
Save
Mark Applied
Hide
Site Reliability Engineer
San Francisco, California, United States
HybridFull Time
Runloop
Runloop: Provides infrastructure and secure sandboxes for AI agents.
5+ YOE5+ years software engineering experience with 3+ years in SRE/DevOps, strong Python or Go skills, containerization, cloud infra, monitoring, networking, Linux administration, on‑call and incident management.
AWS, GCP, Azure, Grafana, Prometheus, Datadog, Python, Go, Docker, Kubernetes, Terraform, Pulumi, Sentry, RUM, CI/CD
1mo
Save
Mark Applied
Hide
Senior Site Reliability Engineer
Sunnyvale, California, United States
$90k-$180k/yr OnsiteFull Time
Abbott
AbbottNYSE: ABT: Manufactures medical devices, diagnostics, and nutritional health products.
Ensure reliability, scalability, and performance of a medical-device remote monitoring platform; expertise in cloud (Azure), Kubernetes, observability, automation, and incident management; bachelor's in a technical discipline.
Python, Go, Bash, PowerShell, Microsoft Azure, Azure Kubernetes Service (AKS), Azure Monitor, Azure DevOps, Azure Policy, Kubernetes, Docker, Prometheus, Grafana, ELK/EFK, Datadog, Linux
2mo
Save
Mark Applied
Hide
Senior Site Reliability Engineer
Oakland, California, United States
$175k-$210k/yr HybridFull Time
Fivetran
Fivetran: Automates data movement into cloud data warehouses.
5+ YOE5+ years SaaS experience; managed Kubernetes, cloud platforms (AWS/GCP/Azure), Terraform/Ansible/ArgoCD; Python/Shell scripting, Linux admin, PostgreSQL; incident response and reliability engineering experience.
Kubernetes, EKS, AKS, GKE, PostgreSQL, ArgoCD, Terraform, Ansible, Python, Shell, Go, Java, AWS, GCP, Azure, Grafana, Buildkite, Temporal, Pulumi, Linux, VPN, PrivateLink, Private Service Connect (GCP)
2d
Save
Mark Applied
Hide
Site Reliability Engineering Manager, Vehicle Software
Sunnyvale, California, United States
$276k-$294k/yr HybridFull Time
Wayve
Wayve: Develops end-to-end artificial intelligence for autonomous driving systems.
8+ YOE3+ MgmtRequires 8+ years building production software systems, 3+ years of people leadership, SRE and reliability expertise, architecture experience, and hands-on coding in C++, Rust, Python, or Go.
C++, Rust, Python, Go, Linux, CI/CD
2d
Save
Mark Applied
Hide
Site Reliability Engineering Manager, Vehicle Software
Sunnyvale, California, United States
$276k-$294k/yr HybridFull Time
Wayve
Wayve: Develops AI software for autonomous vehicle navigation.
8+ YOE3+ MgmtRequires 8+ years building production software systems, 3+ years people leadership, SRE and reliability practices, architecture experience, and production coding in C++, Rust, Python, or Go.
Linux, C++, Rust, Python, Go, CI/CD
1mo
Save
Mark Applied
Hide
Senior Site Reliability Engineer
Sunnyvale or Sylmar
$90k-$180k/yr OnsiteFull Time
Abbott
AbbottNYSE: ABT: Provides medical devices, diagnostics, and science-based nutritional products.
Senior SRE with strong distributed systems, cloud (Azure), Kubernetes, observability, automation, incident management, and cross-functional communication skills for a medical device remote monitoring platform.
Python, Go, Bash, PowerShell, Microsoft Azure, Azure Kubernetes Service (AKS), Azure Monitor, Azure DevOps, Azure Policy, Kubernetes, Docker, Prometheus, Grafana, ELK, EFK, Datadog, Linux
1mo
Save
Mark Applied
Hide
Principal Site Reliability Engineer
San Francisco or Toronto
OnsiteFull Time
Cerebras Systems
Cerebras SystemsNasdaq: CBRS: Manufactures specialized computer chips designed for AI.
15+ YOE15+ years in SRE/infrastructure/platform engineering with large-scale fleets; experience in capacity management, orchestration, observability, SLOs/SLIs, incident response, and cross-team architecture.
Wafer-Scale Engine (WSE), Bazel
2mo
Save
Mark Applied
Hide
Staff Site Reliability Engineer
Foster City, California, United States
$250k-$300k/yr HybridFull Time
Zoox
ZooxNASDAQ: AMZN: Developing autonomous robotaxis for urban ride-hailing services.
5+ YOE5+ years operating GitHub Enterprise at scale, monorepo management, CI/CD integration, infrastructure-as-code (Terraform/Pulumi), cloud platform experience, technical leadership and migration planning.
Git, GitHub Enterprise, GitHub Cloud, Buildkite, GitHub Actions, Jenkins, GitLab CI, Terraform, Pulumi, Bazel, Buck, Reviewable, Gerrit
2mo
Save
Mark Applied
Hide
Site Reliability/Devops Engineer
San Francisco, California, United States
$100k-$200k/yr OnsiteFull Time
Graphon
Graphon: Developing graph-native AI models for multimodal data reasoning.
Proficient in Bash and Python; experience with infrastructure-as-code, Docker, CI/CD, multi-cloud deployments, networking and identity access; comfortable managing production environments and using AI tools.
Bash, Python, Infrastructure-as-code, Docker, CI/CD, AI tools
1mo
Save
Mark Applied
Hide
Senior Site Reliability Engineer
San Francisco, California, United States
$149k-$224k/yr HybridFull Time
Salesforce
SalesforceNYSE: CRM: Sells cloud-based customer relationship management and business software solutions.
5+ YOE5+ years systems and software engineering experience for large-scale internet services; expertise in SRE principles, containers, observability, incident management, Python and Go, and applying AI/ML to operations.
Temporal, Airflow, Argo Workflows, Docker, Kubernetes, DNS, HTTP, Grafana, Prometheus, ELK, Splunk, Datadog, Python, Go, Linux, Claude Code, GitHub Copilot, Codex, Cursor, AWS, GCP, MCP
3mo
Save
Mark Applied
Hide
Staff Site Reliability Engineer
Mountain View, California, United States
$252k-$308k/yr HybridFull Time
EarnIn
EarnIn: Provides immediate access to earned wages through a mobile app.
7+ YOE7+ years in SRE or related field; experience applying AI/LLMs to operations; strong SLO/SLI and incident management; software engineering in Python or Go; observability and IaC proficiency; AI-assisted development tools; fintech/regulated environment experience.
Datadog, CloudWatch, OpenTelemetry, Terraform, Kubernetes, AWS, Python, Go, Cursor, Claude Code, Copilot
2mo
Save
Mark Applied
Hide
Sr. Site Reliability Engineer
Sunnyvale, California, United States
$170k-$196k/yr OnsiteFull Time
Illumio
Illumio: Provides zero-trust segmentation software to contain cyberattacks.
5+ YOE5+ years SRE experience with AWS and/or Azure, scripting in PowerShell/Python/Go, CI/CD experience (Azure DevOps, Jenkins, GitLab CI/CD), containerization knowledge (Docker, Kubernetes), bachelor’s degree or equivalent, on-call and incident management experience.
AWS, Azure, PowerShell, Python, Go, Azure DevOps, Jenkins, GitLab CI/CD, Docker, Kubernetes
2mo
Save
Mark Applied
Hide
Senior Site Reliability Engineer
San Francisco, California, United States
$117k-$209k/yr OnsiteFull Time
Autodesk
AutodeskNASDAQ: ADSK: Developing software for architecture, engineering, and entertainment industries.
7+ YOEU.S. citizen required. 7+ years SRE/platform/cloud experience; B.S. in CS/Engineering or equivalent; experience with large-scale cloud production systems, SLOs/SLIs, observability, incident management, automation, and IaC. Programming in Python/Go/Java/PowerShell/Bash.
AWS, Azure, Python, Go, Java, PowerShell, Bash, Kubernetes, Splunk, Dynatrace, Datadog, CloudWatch, CI/CD
1mo
Save
Mark Applied
Hide
Senior Lead Site Reliability Engineer
Palo Alto, California, United States
$171k-$260k/yr OnsiteFull Time
JPMorgan Chase
JPMorgan ChaseNYSE: JPM: Global financial services firm providing banking and investment solutions.
5+ YOE5+ years applied SRE experience, formal SRE training/certification, expertise in observability, distributed systems, cloud-native and AI-assisted reliability workflows.
Grafana, Dynatrace, Prometheus, Datadog, Splunk, Java, Go (Golang), Python, Terraform, LangChain, LangGraph, AutoGen, CrewAI, GitHub Copilot, Claude, Fluentd, Logstash, Vector, Kafka, RabbitMQ, SQS, Neo4j, TigerGraph, Pinecone, Weaviate, Chroma, Docker, Kubernetes, GitOps, MCP (Model Context Protocol), TensorFlow, PyTorch, scikit-learn, Hadoop, Spark, Flink, MongoDB, Cassandra, DynamoDB, InfluxDB, TimescaleDB, Chaos Monkey, Gremlin, LitmusChaos
2w
Save
Mark Applied
Hide
Sr. Site Reliability Engineer - Paze
Scottsdale or Chicago or San Francisco or New York City or Phoenix or California or Illinois
$106k-$156k/yr HybridFull Time
Early Warning Services
Early Warning Services: Operates payment and risk solutions for the financial industry.
3+ YOEBachelor's degree in business, computer science, or related field; 3+ years of related technical or software development experience; Linux administration, Git, scripting, observability, incident management, and enterprise-scale experience required.
Linux, Git, Java, Ruby, Python, JavaScript, Go, AWS, Docker, Kubernetes, Swarm, CI/CD, TCP/UDP/IP
5d
Save
Mark Applied
Hide
Staff Site Reliability Engineer, Waymo Fleet
San Francisco or Mountain View or United States or North America
$251k-$310k/yr OnsiteFull Time
Waymo
Waymo: Autonomous driving technology for ride-hailing and logistics.
8+ YOERequires 8+ years architecting mission-critical systems in C++, Java, or Python; reliability leadership; cross-functional influence; and a bachelor's degree or 10+ years of similar experience. Advanced degree preferred.
C++, Java, Python
3w
Save
Mark Applied
Hide
Senior Site Reliability Engineer - Infra Ops
San Francisco, California, United States
$153k-$205k/yr RemoteFull Time
Circle
CircleNYSE: CRCL: Digital currency issuer and blockchain financial infrastructure provider.
5+ YOE5+ years SRE/DevOps experience; deep Kubernetes and Terraform expertise; production software development in Go, Python, or JavaScript/TypeScript; cloud networking, observability, SLO/SLI, and on‑call incident management.
Kubernetes, Terraform, Go, Python, JavaScript, TypeScript, CI/CD, GitOps

Explore Jobs

Expand Your Job Search