306 site reliability engineer jobs at 154 companies in California

1w
Save
Mark Applied
Hide
Site Reliability Engineer Intern (Data Infra) - 2027 Fall
San Jose, California, United States
OnsiteInternship
ByteDance
ByteDance: Global technology specializing in AI-powered content platforms.
Currently pursuing a bachelor's degree in computer science or related technical discipline; programming experience in C, C++, Java, Python, Go, or Rust; knowledge of Unix/Linux internals, networking, and distributed systems.
C, C++, Java, Python, Go, Rust, Unix, Linux, Kubernetes, Redis, MySQL, Flink, Nginx, Docker, OpenStack, Hadoop, Spark
3mo
Save
Mark Applied
Hide
Site Reliability Engineer
San Francisco or South San Francisco
$150k/yr OnsiteFull Time
VantageScore Solutions, LLC
VantageScore Solutions, LLC: A Higher Level of Confidence
5+ YOEExperienced Site Reliability Engineer with a DevSecOps focus; patch management, vulnerability remediation; AWS and CI/CD, security tooling.
AWS, EC2, ECS, Lambda, EKS, S3, RDS, IAM, VPC, CloudTrail, Config, GuardDuty, GitHub Actions, CodePipeline, Terraform, CloudFormation, AWS CDK, Kubernetes, Snyk, Wiz, Prisma Cloud, Kong, HashiCorp Vault, Secrets Manager, CloudWatch, Datadog, Grafana
2w
Save
Mark Applied
Hide
Site Reliability Engineer (SRE)
Santa Clara, California, United States
$50-$60/hr RemoteContract
ServiceNow
ServiceNowNYSE: NOW: Enterprise software providing cloud-based workflow automation platforms.
3+ YOEBachelor's degree in computer science or related field; 3+ years in site reliability engineering; 2+ years with AWS and cloud automation; Kubernetes, Linux, Terraform, networking, GitOps, monitoring, and customer support experience.
AWS, Kubernetes, Helm, Linux, Terraform, GitOps, Prometheus, Grafana, Bazel, CueLang, Version Control, Okta, Snowflake, Google
3w
Save
Mark Applied
Hide
Site Reliability Engineer (SRE)
San Francisco, California, United States
$350k-$475k/yr OnsiteFull Time
Thinking Machines Lab
Thinking Machines Lab: Private AI research and product building customizable multimodal systems for researchers and the wider public.
Experience in distributed systems/cloud/site reliability, software automation for reliability, incident response and postmortems, strong communication and coordination skills.
Tinker, Kubernetes, LoRA, CI/CD
2mo
Save
Mark Applied
Hide
Senior Site Reliability Engineer
New York City or Austin or Berlin or Bucharest or Chicago or Dubai or Jakarta or London or Paris or San Francisco or São Paulo or Singapore or Seoul or Sydney or Tokyo
HybridFull Time
Braze
BrazeNASDAQ: BRZE: Customer engagement platform for cross-channel marketing and analytics.
3+ YOE3+ years as a Software/DevOps/Site Reliability Engineer, strong Linux/Unix shell skills, programming experience in Ruby and/or Go, experience with Docker, Kubernetes, Terraform/Chef, and data stores like MongoDB, Redis, Kafka, or Postgres.
Ruby on Rails, Ruby, Go, Linux, Unix Shell, Docker, Kubernetes, Terraform, Chef, MongoDB, Redis, Kafka, Postgres, PagerDuty
3mo
Save
Mark Applied
Hide
Site Reliability Engineer
Europe or Canada or Bellevue or Los Angeles
RemoteFull Time
Shopify
ShopifyNasdaq: SHOP: Provides internet infrastructure and tools for commerce.
Experienced SRE/engineer with on-call experience, ability to build resilient production tooling, improve observability, respond to alerts, and collaborate across engineering teams.
IDE
1mo
Save
Mark Applied
Hide
Site Reliability Engineer
California, United States
OnsiteFull Time
Arena Intelligence
Arena Intelligence: AI model evaluation platform serving enterprises, AI labs, and independent researchers in real-world workflows.
6+ YOE6+ years backend engineering with distributed systems, proficiency in Go or Rust, experience with LLM provider APIs, cloud (AWS/GCP), Kubernetes, Terraform, Postgres, and Redis.
Go, Rust, OpenAI, Anthropic, Google, AWS, GCP, Kubernetes, Terraform, Postgres, Redis, Bifrost, Kong, Envoy, Tyk, Stripe, Metronome, Orb, vLLM, LiteLLM, LangChain
2mo
Save
Mark Applied
Hide
Site Reliability Engineer
California or United States
RemoteFull Time
STN Incorporated
STN Incorporated: U.S.-based IT infrastructure provider delivering managed cloud, cybersecurity, and GPU compute services to enterprises and AI teams.
5+ YOE5+ years in SRE/DevOps or production engineering; strong Go and/or Python skills; Kubernetes at scale; observability with Prometheus, Grafana, Datadog, OpenTelemetry; incident management and on-call experience.
Go, Python, Kubernetes, Prometheus, Grafana, Datadog, OpenTelemetry, Gremlin, Litmus
1mo
Save
Mark Applied
Hide
Site Reliability Engineer
San Francisco or Alpharetta or Arlington or Augusta or Ashburn or Allentown or Appleton or Atlanta or Annapolis Junction or Ann Arbor or Herndon or Allen
$165k-$241k/yr RemoteFull Time
Cisco
CiscoNASDAQ: CSCO: Global leader in networking, cybersecurity, and cloud-native technology solutions.
7+ YOE7+ years SRE or related experience; BS/MS/PhD with corresponding years; U.S. Person required for FedRAMP/IL-5 work; on-call participation; strong coding, automation, reliability, and security skills.
3w
Save
Mark Applied
Hide
Principal Site Reliability Engineer (Hybrid)
Merrimack or San Diego
$118k-$201k/yr HybridFull Time
BAE Systems
BAE SystemsLSE: BA.: Global defense, aerospace, and security technology.
4+ YOERequires 4–6+ years of site reliability engineering, Juniper networking, cloud technologies, automation, storage, virtualization, and security clearance eligibility; Security+ required or obtainable within 90 days.
Juniper, Ansible, Helm Charts, NFS, JDFS, Ceph, S3, VMware, Open Stack, Azure Stack, Kubernetes, Terraform
1mo
Save
Mark Applied
Hide
Site Reliability Engineer
San Francisco, California, United States
HybridFull Time
Runloop AI
Runloop AI: Runloop AI provides AI infrastructure, secure code sandboxes, and evaluation tools for developers building software-engineering agents.
5+ YOE5+ years software engineering experience with 3+ years in SRE/DevOps, strong Python or Go skills, containerization, cloud infra, monitoring, networking, Linux administration, on‑call and incident management.
AWS, GCP, Azure, Grafana, Prometheus, Datadog, Python, Go, Docker, Kubernetes, Terraform, Pulumi, Sentry, RUM, CI/CD
3w
Save
Mark Applied
Hide
Site Reliability Engineer
United States or California
$110k-$253k/yr RemoteFull Time
Veeam Government Solutions
Veeam Government Solutions: The Data and AI Trust.
7+ YOE7+ years software engineering with 3+ years SRE/platform experience, government/regulated cloud familiarity, Azure and IaC experience, programming (TypeScript/JS, Go, Java, C#), observability and CI/CD experience.
Microsoft TFS, Azure DevOps, Git, BitBucket, Azure, Azure Government, Azure Monitor, Application Insights, Cosmos Db, Azure Functions, ARM templates, AWS CloudFormation, Terraform, Terragrunt, Pulumi, Serverless Framework, Kubernetes, Prometheus, Grafana, OpenTelemetry, Elastic Stack, GitHub Actions, GitLab CI, ArgoCD, FluxCD, Dagger, TypeScript, JavaScript, Go, Java, C#
1mo
Save
Mark Applied
Hide
Lead Site Reliability Engineer
Los Angeles, California, United States
$140k-$199k/yr HybridFull Time
Green Dot Corporation
Green Dot CorporationNYSE: GDOT: Public U.S. fintech bank holding providing banking and payment services to consumers and businesses.
7+ YOE7+ years in release/reliability engineering, cloud platform experience (AWS/Azure/GCP), automated deployment and observability proficiency, scripting with PowerShell/Bash/Python, excellent troubleshooting and communication skills.
AWS, Azure, GCP, PowerShell, Bash, Python
2mo
Save
Mark Applied
Hide
Senior Site Reliability Engineer
Oakland, California, United States
$175k-$210k/yr HybridFull Time
Fivetran
Fivetran: Automated data movement and integration platform for organizations.
5+ YOE5+ years SaaS experience; managed Kubernetes, cloud platforms (AWS/GCP/Azure), Terraform/Ansible/ArgoCD; Python/Shell scripting, Linux admin, PostgreSQL; incident response and reliability engineering experience.
Kubernetes, EKS, AKS, GKE, PostgreSQL, ArgoCD, Terraform, Ansible, Python, Shell, Go, Java, AWS, GCP, Azure, Grafana, Buildkite, Temporal, Pulumi, Linux, VPN, PrivateLink, Private Service Connect (GCP)
3mo
Save
Mark Applied
Hide
Founding Engineer - Site Reliability
San Francisco or United States
$185k-$285k/yr RemoteFull Time
uRun
uRun: AI infrastructure helping model labs, builders, and research teams run real-time interactive video and stateful inference.
7+ YOE7+ years in site reliability or infrastructure engineering; strong SLOs, incident response, and observability; Kubernetes and cloud (AWS); software engineering fundamentals; first SRE at a company.
Kubernetes, AWS, Prometheus, Grafana, Datadog, Automation, VPC, GPU compute
1mo
Save
Mark Applied
Hide
Site Reliability Engineer
San Francisco, California, United States
OnsiteFull Time
Specter
Specter: Private San Francis building AI-powered video sensors and wireless networks for industrial businesses.
Strong Linux administration, experience with edge/on‑prem hardware and cloud (AWS), networking fundamentals, scripting in Python/Go/Bash, containerization (Docker, Kubernetes) and embedded/firmware familiarity; on‑call participation.
AWS, Bash, C, Docker, Go, Kubernetes, Linux, Python, Rust, SSH, DNS, VPN, IAM
1mo
Save
Mark Applied
Hide
Site Reliability Engineer
El Segundo, California, United States
$170k-$195k/yr OnsiteFull Time
Picogrid
Picogrid: Private defense technology integrating sensors, platforms, and operators for military, public-safety, and industrial missions.
3+ YOE3+ years SRE experience, deep Kubernetes and Terraform/OpenTofu skills, AWS proficiency, observability (Grafana, Prometheus, Loki, OpenTelemetry), incident response, HA databases, and IoT/edge fleet experience.
Grafana, Prometheus, Loki, OpenTelemetry, Terraform, OpenTofu, Kubernetes, AWS, Nebula, WireGuard, Tailscale, Sloth, Pyrra, NVIDIA Jetson
1mo
Save
Mark Applied
Hide
Site Reliability Engineer
Santa Clara, California, United States
$230k-$250k/yr OnsiteFull Time
Forward Networks
Forward Networks: Private network-software helping enterprises and government agencies model, secure, and automate complex hybrid networks.
6+ YOE6+ years SRE/DevOps experience in SaaS/cloud, strong networking fundamentals, Kubernetes, observability (Prometheus/Grafana/Datadog/Splunk), Python/Bash automation, cloud and IaC (AWS/GCP/Azure, Terraform/Ansible), and incident response ownership.
Kubernetes, Prometheus, Grafana, Datadog, Splunk, Python, Bash, AWS, GCP, Azure, Terraform, Ansible
1mo
Save
Mark Applied
Hide
Site Reliability Engineer
Palo Alto or Newport Beach
$165k-$190k/yr OnsiteFull Time
Obsidian Security
Obsidian Security: Cybersecurity software securing enterprises' SaaS applications, identities, data, and AI agents across third-party applications.
3+ YOE3+ years DevOps/SRE experience on GCP and/or AWS, Bachelor's in CS or related, proficiency with Kubernetes, Helm, GitLab CI/CD, ArgoCD, Prometheus, Grafana; programming in Golang or Python; strong communication and critical thinking.
Kubernetes, Helm, GitLab CI/CD, ArgoCD, Prometheus, Grafana, Golang, Python, Kafka, Elasticsearch, PostgreSQL, ScyllaDB, Databricks, Dagster, Sentry, Kong, AWS, GCP
1mo
Save
Mark Applied
Hide
Site Reliability Engineer (Raptor)
Hawthorne, California, United States
$125k-$175k/yr OnsiteFull Time
SpaceX
SpaceXNasdaq: SPCX: Designing, manufacturing, and launching advanced rockets and spacecraft.
1+ YOE1+ years hands-on experience with client/server hardware, networking, Linux/Windows, scripting and automation; bachelor's in CS/engineering/math or 2+ years software experience in lieu; HPC and systems engineering experience preferred.
Infiniband, ANSYS, StarCCM+, Bash, Python, Puppet, Ansible, Kubernetes, Docker, Linux, Windows

Explore Jobs

Expand Your Job Search