279 site reliablity engineer jobs at 130 companies in California
1w
Save
Mark Applied
Hide
1w
Site Reliability Engineer Intern (Data Infra) - 2027 Fall
San Jose, California, United States
OnsiteInternship
ByteDance: Global technology specializing in AI-powered content platforms.
Currently pursuing a bachelor's degree in computer science or related technical discipline; programming experience in C, C++, Java, Python, Go, or Rust; knowledge of Unix/Linux internals, networking, and distributed systems.
Member of Technical Staff, Site Reliablity Engineer
San Francisco, California, United States
$200k-$270k/yrHybridFull Time
Vapi: Private voice AI platform that lets developers and enterprises build, deploy, and manage conversational voice agents.
Experience running incident command and postmortems, operating SLOs/error budgets, capacity planning and load testing, Kubernetes production ops, KEDA autoscaling, and shipping services in Go or TypeScript.
3+ YOEBachelor's degree in computer science or related field; 3+ years in site reliability engineering; 2+ years with AWS and cloud automation; Kubernetes, Linux, Terraform, networking, GitOps, monitoring, and customer support experience.
AWS, Kubernetes, Helm, Linux, Terraform, GitOps, Prometheus, Grafana, Bazel, CueLang, Version Control, Okta, Snowflake, Google
Thinking Machines Lab: Private AI research and product building customizable multimodal systems for researchers and the wider public.
Experience in distributed systems/cloud/site reliability, software automation for reliability, incident response and postmortems, strong communication and coordination skills.
New York City or Austin or Berlin or Bucharest or Chicago or Dubai or Jakarta or London or Paris or San Francisco or São Paulo or Singapore or Seoul or Sydney or Tokyo
HybridFull Time
BrazeNASDAQ: BRZE: Customer engagement platform for cross-channel marketing and analytics.
3+ YOE3+ years as a Software/DevOps/Site Reliability Engineer, strong Linux/Unix shell skills, programming experience in Ruby and/or Go, experience with Docker, Kubernetes, Terraform/Chef, and data stores like MongoDB, Redis, Kafka, or Postgres.
BAE SystemsLSE: BA.: Global defense, aerospace, and security technology.
4+ YOERequires 4–6+ years of site reliability engineering, Juniper networking, cloud technologies, automation, storage, virtualization, and security clearance eligibility; Security+ required or obtainable within 90 days.
Green Dot CorporationNYSE: GDOT: Public U.S. fintech bank holding providing banking and payment services to consumers and businesses.
7+ YOE7+ years in release/reliability engineering, cloud platform experience (AWS/Azure/GCP), automated deployment and observability proficiency, scripting with PowerShell/Bash/Python, excellent troubleshooting and communication skills.
Arena Intelligence: AI model evaluation platform serving enterprises, AI labs, and independent researchers in real-world workflows.
6+ YOE6+ years backend engineering with distributed systems, proficiency in Go or Rust, experience with LLM provider APIs, cloud (AWS/GCP), Kubernetes, Terraform, Postgres, and Redis.
STN Incorporated: U.S.-based IT infrastructure provider delivering managed cloud, cybersecurity, and GPU compute services to enterprises and AI teams.
5+ YOE5+ years in SRE/DevOps or production engineering; strong Go and/or Python skills; Kubernetes at scale; observability with Prometheus, Grafana, Datadog, OpenTelemetry; incident management and on-call experience.
San Francisco or Alpharetta or Arlington or Augusta or Ashburn or Allentown or Appleton or Atlanta or Annapolis Junction or Ann Arbor or Herndon or Allen
$165k-$241k/yrRemoteFull Time
CiscoNASDAQ: CSCO: Global leader in networking, cybersecurity, and cloud-native technology solutions.
7+ YOE7+ years SRE or related experience; BS/MS/PhD with corresponding years; U.S. Person required for FedRAMP/IL-5 work; on-call participation; strong coding, automation, reliability, and security skills.
Runloop AI: Runloop AI provides AI infrastructure, secure code sandboxes, and evaluation tools for developers building software-engineering agents.
5+ YOE5+ years software engineering experience with 3+ years in SRE/DevOps, strong Python or Go skills, containerization, cloud infra, monitoring, networking, Linux administration, on‑call and incident management.
SpaceXNasdaq: SPCX: Designing, manufacturing, and launching advanced rockets and spacecraft.
1+ YOE1+ years hands-on experience with client/server hardware, networking, Linux/Windows, scripting and automation; bachelor's in CS/engineering/math or 2+ years software experience in lieu; HPC and systems engineering experience preferred.
RobloxNYSE: RBLX: Global platform for user-created immersive digital experiences.
6+ YOE6+ years SRE or software engineering experience; Bachelor’s in Computer Science or equivalent; fluency in Go, Java, or C#; experience with Kubernetes, Nomad, Vault, and Consul; strong reliability and observability practices.
AbbottNYSE: ABT: Global healthcare technology focused on life-changing medical innovations.
Ensure reliability, scalability, and performance of a medical-device remote monitoring platform; expertise in cloud (Azure), Kubernetes, observability, automation, and incident management; bachelor's in a technical discipline.
Python, Go, Bash, PowerShell, Microsoft Azure, Azure Kubernetes Service (AKS), Azure Monitor, Azure DevOps, Azure Policy, Kubernetes, Docker, Prometheus, Grafana, ELK/EFK, Datadog, Linux
K2 Space: Building high-power, mass-produced satellite platforms.
5+ YOE5+ years SRE/DevOps experience or BS in CS/IT/STEM, deep cloud (AWS/GCP/Azure), IaC, Kubernetes, Linux, programming (Go/Python), strong security and reliability experience.
NVIDIANASDAQ: NVDA: Computing platform for AI and accelerated graphics.
8+ YOE8+ years supporting live-site production environments; BS/MS or equivalent; strong Kubernetes, AWS, Python; Akamai/CDN and SRE on-call experience required.
Cerebras SystemsNasdaq Global Select Market: CBRS: Designs processors and systems for AI training and inference.
15+ YOE15+ years in SRE/infrastructure/platform engineering with large-scale fleets; experience in capacity management, orchestration, observability, SLOs/SLIs, incident response, and cross-team architecture.