486 site reliability jobs at 264 companies in California

2mo
Save
Mark Applied
Hide
Site Reliability Engineer
San Francisco or South San Francisco
$150k/yr OnsiteFull Time
VantageScore
VantageScore: Provides credit scoring and data analytics solutions.
5+ YOEExperienced Site Reliability Engineer with a DevSecOps focus; patch management, vulnerability remediation; AWS and CI/CD, security tooling.
AWS, EC2, ECS, Lambda, EKS, S3, RDS, IAM, VPC, CloudTrail, Config, GuardDuty, GitHub Actions, CodePipeline, Terraform, CloudFormation, AWS CDK, Kubernetes, Snyk, Wiz, Prisma Cloud, Kong, HashiCorp Vault, Secrets Manager, CloudWatch, Datadog, Grafana
1w
Save
Mark Applied
Hide
Site Reliability Engineer (SRE)
San Francisco, California, United States
$350k-$475k/yr OnsiteFull Time
Thinking Machines
Thinking Machines: Building AI systems to extend human will and judgment.
Experience in distributed systems/cloud/site reliability, software automation for reliability, incident response and postmortems, strong communication and coordination skills.
Tinker, Kubernetes, LoRA, CI/CD
2mo
Save
Mark Applied
Hide
Founding Engineer - Site Reliability
San Francisco or United States
$185k-$285k/yr RemoteFull Time
uRun
uRun: Infrastructure cloud for interactive, stateful AI inference.
7+ YOE7+ years in site reliability or infrastructure engineering; strong SLOs, incident response, and observability; Kubernetes and cloud (AWS); software engineering fundamentals; first SRE at a company.
Kubernetes, AWS, Prometheus, Grafana, Datadog, Automation, VPC, GPU compute
3d
Save
Mark Applied
Hide
Sr. Site Reliability Engineer
Washington or Palo Alto
$165k-$230k/yr OnsiteFull Time
SpaceX
SpaceX: Designs and launches advanced rockets and satellite internet constellations.
5+ YOEBachelor’s degree and 5+ years of Linux experience, or 7+ years in software, DevOps, or site reliability engineering; 5+ years with Kubernetes; scripting experience; Top Secret clearance required.
Kubernetes, Linux, Bash, Python, Terraform, Ansible, Bazel, Makefiles, TCP/IP
5d
Save
Mark Applied
Hide
Site Reliability Engineer (SRE)
Santa Clara, California, United States
$50-$60/hr RemoteContract
ServiceNow
ServiceNowNYSE: NOW: Enterprise cloud platform for digital workflow automation.
3+ YOEBachelor's degree in computer science or related field; 3+ years in site reliability engineering; 2+ years with AWS and cloud automation; Kubernetes, Linux, Terraform, networking, GitOps, monitoring, and customer support experience.
AWS, Kubernetes, Helm, Linux, Terraform, GitOps, Prometheus, Grafana, Bazel, CueLang, Version Control, Okta, Snowflake, Google
2mo
Save
Mark Applied
Hide
Senior Site Reliability Engineer
New York City or Austin or Berlin or Bucharest or Chicago or Dubai or Jakarta or London or Paris or San Francisco or São Paulo or Singapore or Seoul or Sydney or Tokyo
HybridFull Time
Braze
BrazeNASDAQ: BRZE: Platform for personalized customer engagement and cross-channel messaging.
3+ YOE3+ years as a Software/DevOps/Site Reliability Engineer, strong Linux/Unix shell skills, programming experience in Ruby and/or Go, experience with Docker, Kubernetes, Terraform/Chef, and data stores like MongoDB, Redis, Kafka, or Postgres.
Ruby on Rails, Ruby, Go, Linux, Unix Shell, Docker, Kubernetes, Terraform, Chef, MongoDB, Redis, Kafka, Postgres, PagerDuty
3w
Save
Mark Applied
Hide
Site Reliability Engineer
San Francisco or Alpharetta or Arlington or Augusta or Ashburn or Allentown or Appleton or Atlanta or Annapolis Junction or Ann Arbor or Herndon or Allen
$165k-$241k/yr RemoteFull Time
Cisco
CiscoNASDAQ: CSCO: Develops and sells networking hardware and cybersecurity software.
7+ YOE7+ years SRE or related experience; BS/MS/PhD with corresponding years; U.S. Person required for FedRAMP/IL-5 work; on-call participation; strong coding, automation, reliability, and security skills.
1w
Save
Mark Applied
Hide
Principal Site Reliability Engineer (Hybrid)
Merrimack or San Diego
$118k-$201k/yr HybridFull Time
BAE Systems
BAE SystemsLondon Stock Exchange: BA: Provides advanced defense, aerospace, and security technology solutions.
4+ YOERequires 4–6+ years of site reliability engineering, Juniper networking, cloud technologies, automation, storage, virtualization, and security clearance eligibility; Security+ required or obtainable within 90 days.
Juniper, Ansible, Helm Charts, NFS, JDFS, Ceph, S3, VMware, Open Stack, Azure Stack, Kubernetes, Terraform
3mo
Save
Mark Applied
Hide
Site Reliability Engineer II
Los Angeles, California, United States
$130k-$145k/yr OnsiteFull Time
AXS
AXS: Provides digital ticketing and marketing solutions for live events.
4+ YOE4-6 years in site reliability or DevOps; BA/BS preferred but not required; cloud operations, infrastructure as code, containers/orchestration, CI/CD; programming/scripting ability to automate tasks.
Cloud, Containers, Orchestration, Infrastructure as Code, CI/CD, Python, Bash, Go
2mo
Save
Mark Applied
Hide
Site Reliability Engineer, Robotics
Los Angeles, California, United States
$164k-$270k/yr OnsiteFull Time
Hadrian
Hadrian: Building autonomous factories for aerospace and defense manufacturing.
Own the reliability of production robotics systems; Kubernetes, telemetry, and coding in TypeScript, Python, Golang, or C++. Design SLOs/SLIs and implement observability.
Kubernetes, Prometheus, Telegraf, OpenTelemetry, Datadog, ROS2, OPC UA, Kafka, MQTT, RabbitMQ, GitOps, IaC, TypeScript, Python, Golang, C++
4d
Save
Mark Applied
Hide
Site Reliability Engineer
Irvine, California, United States
$90k-$105k/yr OnsiteFull Time
ICEYE
ICEYE: Operates synthetic aperture radar (SAR) satellite constellations.
2+ YOEBachelor's degree or equivalent practical experience, 2+ years with AWS and Kubernetes, monitoring tools, networking, distributed systems, databases, and Python or similar scripting.
Datadog, LGTM, AWS, Kubernetes, Prometheus, Grafana, Loki, Tempo, Mimir, Python, Terraform, GitHub Actions, Jenkins, GitLab CI, Docker, Fivetran, Databricks, Holistics
1mo
Save
Mark Applied
Hide
Site Reliability Engineer
Santa Clara, California, United States
$230k-$250k/yr OnsiteFull Time
Forward Networks
Forward Networks: Provides a digital twin platform for enterprise network management.
6+ YOE6+ years SRE/DevOps experience in SaaS/cloud, strong networking fundamentals, Kubernetes, observability (Prometheus/Grafana/Datadog/Splunk), Python/Bash automation, cloud and IaC (AWS/GCP/Azure, Terraform/Ansible), and incident response ownership.
Kubernetes, Prometheus, Grafana, Datadog, Splunk, Python, Bash, AWS, GCP, Azure, Terraform, Ansible
2w
Save
Mark Applied
Hide
Site Reliability Engineer
San Mateo or Arizona or California or Colorado or Florida or Georgia or Illinois or Nevada or North Carolina or Oregon or Texas or Utah or Washington
$140k-$150k/yr RemoteFull Time
VyncaCare
VyncaCare: Offers palliative care services and advance care planning technology.
3+ YOE3+ years SRE/DevOps experience, strong AWS and Terraform skills, Kubernetes and Helm experience, observability and incident response knowledge, bachelor's or equivalent, on-call participation, East Coast hours.
AWS, Terraform, Kubernetes, Helm, Prometheus, Grafana, Datadog, CloudWatch, SigNoz, OpenTelemetry, ArgoCD, Flux, PostgreSQL, MySQL, Redshift, ClickHouse, AWS Secrets Manager, HashiCorp Vault, Snowflake, Python, Go, Linux
1mo
Save
Mark Applied
Hide
Site Reliability Engineer
California or United States
RemoteFull Time
STN
STN: Provides high-performance GPU infrastructure, cloud, and managed IT services.
5+ YOE5+ years in SRE/DevOps or production engineering; strong Go and/or Python skills; Kubernetes at scale; observability with Prometheus, Grafana, Datadog, OpenTelemetry; incident management and on-call experience.
Go, Python, Kubernetes, Prometheus, Grafana, Datadog, OpenTelemetry, Gremlin, Litmus
3mo
Save
Mark Applied
Hide
Site Reliability Engineering
Foster City, California, United States
$140k-$230k/yr HybridFull Time
Zoox
ZooxNASDAQ: AMZN: Developing autonomous robotaxis for urban ride-hailing services.
5+ YOE5+ years SRE/Distributed systems; cloud platforms (AWS, GCP, or Azure); IaC (Terraform, Ansible, Salt, CloudFormation); Kubernetes; Python/Go/C/C++/Java.
AWS, GCP, Azure, Terraform, Ansible, Salt, CloudFormation, Kubernetes, Python, Go, C/C++, Java
1mo
Save
Mark Applied
Hide
Site Reliability Engineer
San Francisco, California, United States
HybridFull Time
Runloop
Runloop: Provides infrastructure and secure sandboxes for AI agents.
5+ YOE5+ years software engineering experience with 3+ years in SRE/DevOps, strong Python or Go skills, containerization, cloud infra, monitoring, networking, Linux administration, on‑call and incident management.
AWS, GCP, Azure, Grafana, Prometheus, Datadog, Python, Go, Docker, Kubernetes, Terraform, Pulumi, Sentry, RUM, CI/CD
1mo
Save
Mark Applied
Hide
Site Reliability Engineer
Lincoln or San Francisco
$125k-$165k/yr RemoteFull Time
TELCOR
TELCOR: Provides healthcare software for laboratory and point-of-care operations.
2+ YOEExperience with distributed systems, Redis, queuing systems, Kubernetes (2+ yrs), AWS, Terraform (2+ yrs), observability, and production operations.
Redis, Kubernetes, AWS, Terraform
1mo
Save
Mark Applied
Hide
Site Reliability Engineer
Mountain View, California, United States
$189k-$232k/yr HybridFull Time
EarnIn
EarnIn: Provides immediate access to earned wages through a mobile app.
3+ YOE3+ years SRE or related experience; hands-on Python/Go coding; experience with observability, incident response, SLOs/SLIs, and distributed systems; strong communication and documentation skills.
Python, Go, Datadog, CloudWatch, logs, metrics, traces, APM, GitHub Copilot, Cursor, ChatGPT, Claude
1mo
Save
Mark Applied
Hide
Site Reliability Engineer
Santa Clara or St. Louis or Bangalore or London or Paris or Melbourne or Taipei or Tokyo
OnsiteFull Time
Netskope
NetskopeNASDAQ: NTSK: Cloud-native cybersecurity and data protection platform for enterprises.
3+ YOEBachelor's in CS/Engineering or equivalent; 3+ years building/managing complex systems (including 1-2 years SRE); experience with cloud services, microservices, availability/performance optimization, debugging, and strong communication.
Python, C, C++, Go, Rust, Docker, Kubernetes, AWS, GCP, KVM, OpenNebula, OpenStack, TCP/IP
1mo
Save
Mark Applied
Hide
Site Reliability Engineer
Palo Alto or Newport Beach
$165k-$190k/yr OnsiteFull Time
Obsidian Security
Obsidian Security: Provides cybersecurity and threat detection for enterprise SaaS applications.
3+ YOE3+ years DevOps/SRE experience on GCP and/or AWS, Bachelor's in CS or related, proficiency with Kubernetes, Helm, GitLab CI/CD, ArgoCD, Prometheus, Grafana; programming in Golang or Python; strong communication and critical thinking.
Kubernetes, Helm, GitLab CI/CD, ArgoCD, Prometheus, Grafana, Golang, Python, Kafka, Elasticsearch, PostgreSQL, ScyllaDB, Databricks, Dagster, Sentry, Kong, AWS, GCP