306 site reliability engineer jobs at 154 companies in California
1w
Save
Mark Applied
Hide
1w
Site Reliability Engineer Intern (Data Infra) - 2027 Fall
San Jose, California, United States
OnsiteInternship
ByteDance: Global technology specializing in AI-powered content platforms.
Currently pursuing a bachelor's degree in computer science or related technical discipline; programming experience in C, C++, Java, Python, Go, or Rust; knowledge of Unix/Linux internals, networking, and distributed systems.
3+ YOEBachelor's degree in computer science or related field; 3+ years in site reliability engineering; 2+ years with AWS and cloud automation; Kubernetes, Linux, Terraform, networking, GitOps, monitoring, and customer support experience.
AWS, Kubernetes, Helm, Linux, Terraform, GitOps, Prometheus, Grafana, Bazel, CueLang, Version Control, Okta, Snowflake, Google
Thinking Machines Lab: Private AI research and product building customizable multimodal systems for researchers and the wider public.
Experience in distributed systems/cloud/site reliability, software automation for reliability, incident response and postmortems, strong communication and coordination skills.
New York City or Austin or Berlin or Bucharest or Chicago or Dubai or Jakarta or London or Paris or San Francisco or São Paulo or Singapore or Seoul or Sydney or Tokyo
HybridFull Time
BrazeNASDAQ: BRZE: Customer engagement platform for cross-channel marketing and analytics.
3+ YOE3+ years as a Software/DevOps/Site Reliability Engineer, strong Linux/Unix shell skills, programming experience in Ruby and/or Go, experience with Docker, Kubernetes, Terraform/Chef, and data stores like MongoDB, Redis, Kafka, or Postgres.
ShopifyNasdaq: SHOP: Provides internet infrastructure and tools for commerce.
Experienced SRE/engineer with on-call experience, ability to build resilient production tooling, improve observability, respond to alerts, and collaborate across engineering teams.
Arena Intelligence: AI model evaluation platform serving enterprises, AI labs, and independent researchers in real-world workflows.
6+ YOE6+ years backend engineering with distributed systems, proficiency in Go or Rust, experience with LLM provider APIs, cloud (AWS/GCP), Kubernetes, Terraform, Postgres, and Redis.
STN Incorporated: U.S.-based IT infrastructure provider delivering managed cloud, cybersecurity, and GPU compute services to enterprises and AI teams.
5+ YOE5+ years in SRE/DevOps or production engineering; strong Go and/or Python skills; Kubernetes at scale; observability with Prometheus, Grafana, Datadog, OpenTelemetry; incident management and on-call experience.
San Francisco or Alpharetta or Arlington or Augusta or Ashburn or Allentown or Appleton or Atlanta or Annapolis Junction or Ann Arbor or Herndon or Allen
$165k-$241k/yrRemoteFull Time
CiscoNASDAQ: CSCO: Global leader in networking, cybersecurity, and cloud-native technology solutions.
7+ YOE7+ years SRE or related experience; BS/MS/PhD with corresponding years; U.S. Person required for FedRAMP/IL-5 work; on-call participation; strong coding, automation, reliability, and security skills.
BAE SystemsLSE: BA.: Global defense, aerospace, and security technology.
4+ YOERequires 4–6+ years of site reliability engineering, Juniper networking, cloud technologies, automation, storage, virtualization, and security clearance eligibility; Security+ required or obtainable within 90 days.
Runloop AI: Runloop AI provides AI infrastructure, secure code sandboxes, and evaluation tools for developers building software-engineering agents.
5+ YOE5+ years software engineering experience with 3+ years in SRE/DevOps, strong Python or Go skills, containerization, cloud infra, monitoring, networking, Linux administration, on‑call and incident management.
Green Dot CorporationNYSE: GDOT: Public U.S. fintech bank holding providing banking and payment services to consumers and businesses.
7+ YOE7+ years in release/reliability engineering, cloud platform experience (AWS/Azure/GCP), automated deployment and observability proficiency, scripting with PowerShell/Bash/Python, excellent troubleshooting and communication skills.
uRun: AI infrastructure helping model labs, builders, and research teams run real-time interactive video and stateful inference.
7+ YOE7+ years in site reliability or infrastructure engineering; strong SLOs, incident response, and observability; Kubernetes and cloud (AWS); software engineering fundamentals; first SRE at a company.
Specter: Private San Francis building AI-powered video sensors and wireless networks for industrial businesses.
Strong Linux administration, experience with edge/on‑prem hardware and cloud (AWS), networking fundamentals, scripting in Python/Go/Bash, containerization (Docker, Kubernetes) and embedded/firmware familiarity; on‑call participation.
Picogrid: Private defense technology integrating sensors, platforms, and operators for military, public-safety, and industrial missions.
3+ YOE3+ years SRE experience, deep Kubernetes and Terraform/OpenTofu skills, AWS proficiency, observability (Grafana, Prometheus, Loki, OpenTelemetry), incident response, HA databases, and IoT/edge fleet experience.
Obsidian Security: Cybersecurity software securing enterprises' SaaS applications, identities, data, and AI agents across third-party applications.
3+ YOE3+ years DevOps/SRE experience on GCP and/or AWS, Bachelor's in CS or related, proficiency with Kubernetes, Helm, GitLab CI/CD, ArgoCD, Prometheus, Grafana; programming in Golang or Python; strong communication and critical thinking.
SpaceXNasdaq: SPCX: Designing, manufacturing, and launching advanced rockets and spacecraft.
1+ YOE1+ years hands-on experience with client/server hardware, networking, Linux/Windows, scripting and automation; bachelor's in CS/engineering/math or 2+ years software experience in lieu; HPC and systems engineering experience preferred.