156 site reliablity engineer jobs at 89 companies in Berkeley, CA
2mo
Save
Mark Applied
Hide
2mo
Member of Technical Staff, Site Reliablity Engineer
San Francisco, California, United States
$200k-$270k/yrHybridFull Time
Vapi: Build and deploy AI-powered voice agents via flexible APIs.
Experience running incident command and postmortems, operating SLOs/error budgets, capacity planning and load testing, Kubernetes production ops, KEDA autoscaling, and shipping services in Go or TypeScript.
Thinking Machines: Building AI systems to extend human will and judgment.
Experience in distributed systems/cloud/site reliability, software automation for reliability, incident response and postmortems, strong communication and coordination skills.
New York City or Austin or Berlin or Bucharest or Chicago or Dubai or Jakarta or London or Paris or San Francisco or São Paulo or Singapore or Seoul or Sydney or Tokyo
HybridFull Time
BrazeNASDAQ: BRZE: Platform for personalized customer engagement and cross-channel messaging.
3+ YOE3+ years as a Software/DevOps/Site Reliability Engineer, strong Linux/Unix shell skills, programming experience in Ruby and/or Go, experience with Docker, Kubernetes, Terraform/Chef, and data stores like MongoDB, Redis, Kafka, or Postgres.
Runloop: Provides infrastructure and secure sandboxes for AI agents.
5+ YOE5+ years software engineering experience with 3+ years in SRE/DevOps, strong Python or Go skills, containerization, cloud infra, monitoring, networking, Linux administration, on‑call and incident management.
Santa Clara or St. Louis or Bangalore or London or Paris or Melbourne or Taipei or Tokyo
OnsiteFull Time
NetskopeNASDAQ: NTSK: Cloud-native cybersecurity and data protection platform for enterprises.
3+ YOEBachelor's in CS/Engineering or equivalent; 3+ years building/managing complex systems (including 1-2 years SRE); experience with cloud services, microservices, availability/performance optimization, debugging, and strong communication.
uRun: Infrastructure cloud for interactive, stateful AI inference.
7+ YOE7+ years in site reliability or infrastructure engineering; strong SLOs, incident response, and observability; Kubernetes and cloud (AWS); software engineering fundamentals; first SRE at a company.
RobloxNYSE: RBLX: Platform for creating and playing user-generated 3D digital experiences.
6+ YOE6+ years SRE or software engineering experience; Bachelor’s in Computer Science or equivalent; fluency in Go, Java, or C#; experience with Kubernetes, Nomad, Vault, and Consul; strong reliability and observability practices.
Obsidian Security: Provides cybersecurity and threat detection for enterprise SaaS applications.
5+ YOE5+ years in SRE/Production Engineering; 3+ years in senior/leadership role; expertise in AWS/GCP, Kubernetes/Helm, observability stacks, CI/CD; experience with multi-tenant SaaS reliability.
Site Reliability Engineer - Hardware Infrastructure
Santa Clara, California, United States
$184k-$357k/yrOnsiteFull Time
NVIDIANASDAQ: NVDA: Designs graphics processing units and artificial intelligence hardware.
8+ YOEDegree in CS or related field (or equivalent experience), 8+ years SRE/DevOps/Production Engineering, SRE principles, infrastructure automation, production reliability, Python/Go/Perl/Ruby, Prometheus and Grafana, strong communication.
AbbottNYSE: ABT: Manufactures medical devices, diagnostics, and nutritional health products.
Ensure reliability, scalability, and performance of a medical-device remote monitoring platform; expertise in cloud (Azure), Kubernetes, observability, automation, and incident management; bachelor's in a technical discipline.
Python, Go, Bash, PowerShell, Microsoft Azure, Azure Kubernetes Service (AKS), Azure Monitor, Azure DevOps, Azure Policy, Kubernetes, Docker, Prometheus, Grafana, ELK/EFK, Datadog, Linux
Cerebras SystemsNasdaq: CBRS: Manufactures specialized computer chips designed for AI.
15+ YOE15+ years in SRE/infrastructure/platform engineering with large-scale fleets; experience in capacity management, orchestration, observability, SLOs/SLIs, incident response, and cross-team architecture.
Stuut: Automates business accounts receivable and collections through AI agents.
7+ YOE7+ years in SRE/infrastructure or backend engineering. Experience with AWS, Kubernetes/EKS, Docker, observability, SLOs/SLIs, Python or TypeScript, CI/CD, and production-grade distributed systems.
Ivo: AI-powered contract review and intelligence platform for legal teams.
5+ YOEMinimum 5 years experience; own uptime and reliability, define SLIs/SLOs/SLAs, design failover and disaster recovery, implement security controls, lead incident response; familiarity with cloud and LLM-driven systems.
IXL Learning: Provides personalized digital learning platforms and educational resources.
6+ YOEBachelor's degree,6+ years SRE/software engineering,experience with OO and scripting languages,cloud (AWS/GCP),Docker/Kubernetes,monitoring,on-call availability,strong troubleshooting and communication skills.
SalesforceNYSE: CRM: Sells cloud-based customer relationship management and business software solutions.
5+ YOE5+ years systems and software engineering experience for large-scale internet services; expertise in SRE principles, containers, observability, incident management, Python and Go, and applying AI/ML to operations.
Pleasanton or Austin or San Francisco or United States
OnsiteFull Time
OracleNYSE: ORCL: Provides cloud infrastructure and enterprise software for global businesses.
8+ YOESenior SRE with strong infrastructure, automation, and programming experience (Terraform, Chef, Ansible, Python, Java, Bash). Minimum multi-year experience in software engineering or equivalent; participates in on-call and incident response.