115 site reliability manager jobs at 40 companies in Fairfax, CA

3mo
Save
Mark Applied
Hide
Site Reliability Engineer
San Francisco or South San Francisco
$150k/yr OnsiteFull Time
VantageScore Solutions, LLC
VantageScore Solutions, LLC: A Higher Level of Confidence
5+ YOEExperienced Site Reliability Engineer with a DevSecOps focus; patch management, vulnerability remediation; AWS and CI/CD, security tooling.
AWS, EC2, ECS, Lambda, EKS, S3, RDS, IAM, VPC, CloudTrail, Config, GuardDuty, GitHub Actions, CodePipeline, Terraform, CloudFormation, AWS CDK, Kubernetes, Snyk, Wiz, Prisma Cloud, Kong, HashiCorp Vault, Secrets Manager, CloudWatch, Datadog, Grafana
3d
Save
Mark Applied
Hide
Manager, Site Reliability Engineer
San Francisco or New York City
$150k-$220k/yr OnsiteFull Time
Forge Global
Forge GlobalNYSE: FRGE: Financial technology operating a private-market marketplace and data, custody, and investment solutions for companies and investors.
10+ YOE5+ MgmtRequires 5+ years leading SRE, DevOps, cloud operations, or reliability functions; 10+ years in engineering or operations; bachelor's degree or equivalent; cloud infrastructure, distributed systems, observability, CI/CD, and automation experience.
AWS, Azure, Kubernetes, Terraform, Ansible, Datadog, CloudWatch
11h
Save
Mark Applied
Hide
Staff Site Reliability Engineer, Ads
San Francisco or United States
$217k-$304k/yr RemoteFull Time, Contract
Reddit
RedditNYSE: RDDT: Social news aggregation, web content rating, and discussion platform.
8+ YOE8+ years in site reliability or infrastructure engineering, distributed systems, cloud-native architecture, observability, automation, incident management, performance optimization, and backend software engineering.
Go, Kubernetes, Kafka, ClickHouse, Spark, Flink, BigQuery
1mo
Save
Mark Applied
Hide
Manager, Site Reliability Engineering
San Francisco, California, United States
$204k-$306k/yr HybridFull Time
Okta
OktaNASDAQ: OKTA: Identity management and access control software provider.
3+ Mgmt3+ years technical leadership experience; experience with cloud-native architectures, Kubernetes, Terraform, CI/CD, observability platforms; strong software development and automation background; US Person status required.
Amazon Web Services (AWS), Kubernetes, Terraform, Grafana, Splunk, APM, CI/CD
1w
Save
Mark Applied
Hide
Senior Manager, Site Reliability Engineering
Mountain View or Mountain View or California or United States
$222k-$301k/yr OnsiteFull Time
Intuit
IntuitNASDAQ: INTU: A global financial technology platform powering prosperity.
8+ YOE3+ Mgmt8+ years in systems, SRE, or infrastructure engineering; 3+ years managing engineering teams; AWS at scale; distributed systems, Kubernetes, IaC, observability, incident management, and AI Ops experience; bachelor's degree required.
AWS, Amazon EC2, Amazon EKS, Amazon ECS, Amazon VPC, Amazon RDS, Amazon DynamoDB, AWS IAM, Amazon CloudWatch, AWS Auto Scaling, Kubernetes, Terraform, AWS CloudFormation, Datadog, Splunk, PagerDuty, Prometheus, Grafana, AI Ops, AIOps platforms
1mo
Save
Mark Applied
Hide
Site Reliability Engineer
San Francisco, California, United States
HybridFull Time
Runloop AI
Runloop AI: Runloop AI provides AI infrastructure, secure code sandboxes, and evaluation tools for developers building software-engineering agents.
5+ YOE5+ years software engineering experience with 3+ years in SRE/DevOps, strong Python or Go skills, containerization, cloud infra, monitoring, networking, Linux administration, on‑call and incident management.
AWS, GCP, Azure, Grafana, Prometheus, Datadog, Python, Go, Docker, Kubernetes, Terraform, Pulumi, Sentry, RUM, CI/CD
2mo
Save
Mark Applied
Hide
Senior Site Reliability Engineer
Oakland, California, United States
$175k-$210k/yr HybridFull Time
Fivetran
Fivetran: Automated data movement and integration platform for organizations.
5+ YOE5+ years SaaS experience; managed Kubernetes, cloud platforms (AWS/GCP/Azure), Terraform/Ansible/ArgoCD; Python/Shell scripting, Linux admin, PostgreSQL; incident response and reliability engineering experience.
Kubernetes, EKS, AKS, GKE, PostgreSQL, ArgoCD, Terraform, Ansible, Python, Shell, Go, Java, AWS, GCP, Azure, Grafana, Buildkite, Temporal, Pulumi, Linux, VPN, PrivateLink, Private Service Connect (GCP)
1mo
Save
Mark Applied
Hide
Principal Site Reliability Engineer
San Francisco or Toronto
OnsiteFull Time
Cerebras Systems
Cerebras SystemsNasdaq Global Select Market: CBRS: Designs processors and systems for AI training and inference.
15+ YOE15+ years in SRE/infrastructure/platform engineering with large-scale fleets; experience in capacity management, orchestration, observability, SLOs/SLIs, incident response, and cross-team architecture.
Wafer-Scale Engine (WSE), Bazel
1d
Save
Mark Applied
Hide
Sr. Site Reliability Engineer
Scottsdale or San Francisco or Chicago or New York City or Phoenix or Chicago or San Francisco
$118k-$183k/yr HybridFull Time
Early Warning Services
Early Warning Services: U.S. bank-owned fintech and consumer reporting agency providing identity, fraud-risk, and real-time payment solutions to financial institutions.
5+ YOEBachelor's degree in business, computer science, or related field; 5+ years of technical experience; incident management, Linux, scripting, observability, Git, security protocols, and enterprise production experience.
Observability, Linux, Git, Java, Ruby, Python, JavaScript, Go, CI/CD, TCP/UDP/IP, AWS, Docker, Kubernetes, Swarm
2mo
Save
Mark Applied
Hide
Staff Site Reliability Engineer
Foster City, California, United States
$250k-$300k/yr HybridFull Time
Zoox
Zoox: Autonomous mobility developing a fully electric robotaxi fleet.
5+ YOE5+ years operating GitHub Enterprise at scale, monorepo management, CI/CD integration, infrastructure-as-code (Terraform/Pulumi), cloud platform experience, technical leadership and migration planning.
Git, GitHub Enterprise, GitHub Cloud, Buildkite, GitHub Actions, Jenkins, GitLab CI, Terraform, Pulumi, Bazel, Buck, Reviewable, Gerrit
2mo
Save
Mark Applied
Hide
Site Reliability/Devops Engineer
San Francisco, California, United States
$100k-$200k/yr OnsiteFull Time
Graphon AI
Graphon AI: Graphon AI is a private enterprise-AI software building multimodal relational memory for organizations and AI agents.
Proficient in Bash and Python; experience with infrastructure-as-code, Docker, CI/CD, multi-cloud deployments, networking and identity access; comfortable managing production environments and using AI tools.
Bash, Python, Infrastructure-as-code, Docker, CI/CD, AI tools
1mo
Save
Mark Applied
Hide
Senior Site Reliability Engineer
San Francisco, California, United States
$149k-$224k/yr HybridFull Time
Salesforce
SalesforceNYSE: CRM: The #1 AI CRM driving customer success together.
5+ YOE5+ years systems and software engineering experience for large-scale internet services; expertise in SRE principles, containers, observability, incident management, Python and Go, and applying AI/ML to operations.
Temporal, Airflow, Argo Workflows, Docker, Kubernetes, DNS, HTTP, Grafana, Prometheus, ELK, Splunk, Datadog, Python, Go, Linux, Claude Code, GitHub Copilot, Codex, Cursor, AWS, GCP, MCP
3mo
Save
Mark Applied
Hide
Staff Site Reliability Engineer
Mountain View, California, United States
$252k-$308k/yr HybridFull Time
EarnIn
EarnIn: Fintech helping workers access earned wages in real time and manage finances without interest or mandatory fees.
7+ YOE7+ years in SRE or related field; experience applying AI/LLMs to operations; strong SLO/SLI and incident management; software engineering in Python or Go; observability and IaC proficiency; AI-assisted development tools; fintech/regulated environment experience.
Datadog, CloudWatch, OpenTelemetry, Terraform, Kubernetes, AWS, Python, Go, Cursor, Claude Code, Copilot
2mo
Save
Mark Applied
Hide
Senior Site Reliability Engineer
San Francisco, California, United States
$117k-$209k/yr OnsiteFull Time
Autodesk
AutodeskNASDAQ: ADSK: Global provider of software for design, engineering, and manufacturing.
7+ YOEU.S. citizen required. 7+ years SRE/platform/cloud experience; B.S. in CS/Engineering or equivalent; experience with large-scale cloud production systems, SLOs/SLIs, observability, incident management, automation, and IaC. Programming in Python/Go/Java/PowerShell/Bash.
AWS, Azure, Python, Go, Java, PowerShell, Bash, Kubernetes, Splunk, Dynatrace, Datadog, CloudWatch, CI/CD
1mo
Save
Mark Applied
Hide
Senior Lead Site Reliability Engineer
Palo Alto, California, United States
$171k-$260k/yr OnsiteFull Time
JPMorgan Chase
JPMorgan ChaseNYSE: JPM: Global financial services and investment banking firm.
5+ YOE5+ years applied SRE experience, formal SRE training/certification, expertise in observability, distributed systems, cloud-native and AI-assisted reliability workflows.
Grafana, Dynatrace, Prometheus, Datadog, Splunk, Java, Go (Golang), Python, Terraform, LangChain, LangGraph, AutoGen, CrewAI, GitHub Copilot, Claude, Fluentd, Logstash, Vector, Kafka, RabbitMQ, SQS, Neo4j, TigerGraph, Pinecone, Weaviate, Chroma, Docker, Kubernetes, GitOps, MCP (Model Context Protocol), TensorFlow, PyTorch, scikit-learn, Hadoop, Spark, Flink, MongoDB, Cassandra, DynamoDB, InfluxDB, TimescaleDB, Chaos Monkey, Gremlin, LitmusChaos
1d
Save
Mark Applied
Hide
Lead Site Reliability Engineer
San Francisco or Richmond
$147k-$234k/yr OnsiteFull Time
Federal Reserve Bank of San Francisco
Federal Reserve Bank of San Francisco: The U.S. central banking system sets monetary policy, supervises banks, and provides payment services.
7+ YOE3+ MgmtBachelor's degree or equivalent practical experience, 7+ years in SRE or DevOps, 3+ years in a lead or senior technical role, and expertise in AWS, Java, Python, Node.js, Terraform, GitLab, Docker, Kubernetes, security, and observability.
Java, Python, Node.js, AWS, Terraform, Lambda, ECS, EC2, Fargate, S3, EBS, EFS, RDS, DynamoDB, Aurora, VPC, Route53, CloudFront, API Gateway, CloudWatch, X-Ray, GitLab, Docker, Kubernetes, SAST, DAST, OWASP Top 10, IAM, Grafana, Datadog, New Relic, CloudWatch Logs, Splunk, LLMs, Agile, Scrum
6d
Save
Mark Applied
Hide
Staff Site Reliability Engineer, Waymo Fleet
San Francisco or Mountain View or United States or North America
$251k-$310k/yr OnsiteFull Time
Waymo
Waymo: Autonomous driving technology and robotaxi service provider.
8+ YOERequires 8+ years architecting mission-critical systems in C++, Java, or Python; reliability leadership; cross-functional influence; and a bachelor's degree or 10+ years of similar experience. Advanced degree preferred.
C++, Java, Python
2mo
Save
Mark Applied
Hide
Senior Software Engineer, Site Reliability Engineering
San Francisco or San Jose or New York City or Seattle or Austin or Washington or California or Massachusetts or New Jersey or Washington or United States
$179k-$273k/yr RemoteFull Time
Thumbtack
Thumbtack: Home services marketplace helping homeowners find and hire local professionals for repairs, maintenance, and improvements.
5+ YOE5+ years managing infrastructure and systems; extensive AWS and Linux fluency; proficiency in Python, Go, PHP, and JavaScript; experience with distributed systems, observability, and on-call rotations; strong communication and troubleshooting skills.
AWS, Linux, Python, Go, PHP, JavaScript, DNS, TLS, HTTP/S, TCP/IP
3w
Save
Mark Applied
Hide
Senior Site Reliability Engineer - Infra Ops
San Francisco, California, United States
$153k-$205k/yr RemoteFull Time
Circle Internet Group, Inc.
Circle Internet Group, Inc.NYSE: CRCL: Public financial technology providing stablecoin, digital-asset, payments, and blockchain infrastructure to businesses and developers.
5+ YOE5+ years SRE/DevOps experience; deep Kubernetes and Terraform expertise; production software development in Go, Python, or JavaScript/TypeScript; cloud networking, observability, SLO/SLI, and on‑call incident management.
Kubernetes, Terraform, Go, Python, JavaScript, TypeScript, CI/CD, GitOps
2w
Save
Mark Applied
Hide
Senior Site Reliability Engineer - Managed Kubernetes
San Francisco or San Jose or Bellevue
$240k-$356k/yr HybridFull Time
Lambda
Lambda: AI infrastructure building GPU cloud services and supercomputers for researchers, enterprises, and hyperscalers.
6+ YOERequires 6+ years in SRE or operations, deep Linux and production Kubernetes expertise, strong Go and Python skills, GitOps, Helm, observability, CI/CD, and Kubernetes provisioning experience.
Kubernetes, Python, Golang, GitOps, ArgoCD, Helm, Linux, EKS, GKE, Prometheus, Grafana, FluentBit, CI/CD, kubeadm, Cluster API, CRDs, CSI, CNI

Explore Jobs

Expand Your Job Search