179 site reliability engineering jobs at 95 companies in San Jose, CA
1mo
Save
Mark Applied
Hide
1mo
Manager, Site Reliability Engineering
San Francisco, California, United States
$204k-$306k/yrHybridFull Time
OktaNASDAQ: OKTA: Provide secure identity management and authentication for enterprises.
3+ Mgmt3+ years technical leadership experience; experience with cloud-native architectures, Kubernetes, Terraform, CI/CD, observability platforms; strong software development and automation background; US Person status required.
Amazon Web Services (AWS), Kubernetes, Terraform, Grafana, Splunk, APM, CI/CD
ServiceNowNYSE: NOW: Enterprise cloud platform for digital workflow automation.
3+ YOEBachelor's degree in computer science or related field; 3+ years in site reliability engineering; 2+ years with AWS and cloud automation; Kubernetes, Linux, Terraform, networking, GitOps, monitoring, and customer support experience.
AWS, Kubernetes, Helm, Linux, Terraform, GitOps, Prometheus, Grafana, Bazel, CueLang, Version Control, Okta, Snowflake, Google
Thinking Machines: Building AI systems to extend human will and judgment.
Experience in distributed systems/cloud/site reliability, software automation for reliability, incident response and postmortems, strong communication and coordination skills.
New York City or Austin or Berlin or Bucharest or Chicago or Dubai or Jakarta or London or Paris or San Francisco or São Paulo or Singapore or Seoul or Sydney or Tokyo
HybridFull Time
BrazeNASDAQ: BRZE: Platform for personalized customer engagement and cross-channel messaging.
3+ YOE3+ years as a Software/DevOps/Site Reliability Engineer, strong Linux/Unix shell skills, programming experience in Ruby and/or Go, experience with Docker, Kubernetes, Terraform/Chef, and data stores like MongoDB, Redis, Kafka, or Postgres.
San Francisco or Alpharetta or Arlington or Augusta or Ashburn or Allentown or Appleton or Atlanta or Annapolis Junction or Ann Arbor or Herndon or Allen
$165k-$241k/yrRemoteFull Time
CiscoNASDAQ: CSCO: Develops and sells networking hardware and cybersecurity software.
7+ YOE7+ years SRE or related experience; BS/MS/PhD with corresponding years; U.S. Person required for FedRAMP/IL-5 work; on-call participation; strong coding, automation, reliability, and security skills.
Runloop: Provides infrastructure and secure sandboxes for AI agents.
5+ YOE5+ years software engineering experience with 3+ years in SRE/DevOps, strong Python or Go skills, containerization, cloud infra, monitoring, networking, Linux administration, on‑call and incident management.
Santa Clara or St. Louis or Bangalore or London or Paris or Melbourne or Taipei or Tokyo
OnsiteFull Time
NetskopeNASDAQ: NTSK: Cloud-native cybersecurity and data protection platform for enterprises.
3+ YOEBachelor's in CS/Engineering or equivalent; 3+ years building/managing complex systems (including 1-2 years SRE); experience with cloud services, microservices, availability/performance optimization, debugging, and strong communication.
Director, Engineering Operations and Site Reliability Engineering - Datacenter Server Systems
Santa Clara, California, United States
$292k-$443k/yrOnsiteFull Time
NVIDIANASDAQ: NVDA: Designs GPU-accelerated computing and artificial intelligence hardware.
12+ YOE7+ MgmtBS/MS in CS/EE/CE or equivalent; 12+ years in infrastructure/systems/reliability/datacenter ops, including 7+ years people management; strong Linux, server, networking, monitoring, incident management, automation skills; strong communication.
Site Reliability Engineering (SRE) Manager, Apple Maps
Cupertino, California, United States
OnsiteFull Time
AppleNASDAQ: AAPL: Designs and sells consumer electronics, software, and online services.
Build, manage, and deliver highly available, automated infrastructure for Apple Maps at global scale; focus on reliability, scalability, and operational excellence.
Senior Software Engineer, Site Reliability Engineering
New York or San Ramon or Reno
$153k-$210k/yrHybridFull Time
Ridgeline: Cloud-native platform for investment management operations.
3+ YOE3–6 years SRE/DevOps experience, 2+ years on AWS, proficiency with Terraform, observability, CI/CD, Python/Go/Bash, incident response, and strong communication and troubleshooting skills.
Senior Software Engineer, Site Reliability Engineering
San Francisco or San Jose or New York City or Seattle or Austin or Washington or California or Massachusetts or New Jersey or Washington or United States
$179k-$273k/yrRemoteFull Time
Thumbtack: Online marketplace connecting homeowners with local service professionals.
5+ YOE5+ years managing infrastructure and systems; extensive AWS and Linux fluency; proficiency in Python, Go, PHP, and JavaScript; experience with distributed systems, observability, and on-call rotations; strong communication and troubleshooting skills.
uRun: Infrastructure cloud for interactive, stateful AI inference.
7+ YOE7+ years in site reliability or infrastructure engineering; strong SLOs, incident response, and observability; Kubernetes and cloud (AWS); software engineering fundamentals; first SRE at a company.
San Mateo or Arizona or California or Colorado or Florida or Georgia or Illinois or Nevada or North Carolina or Oregon or Texas or Utah or Washington
$140k-$150k/yrRemoteFull Time
VyncaCare: Offers palliative care services and advance care planning technology.
3+ YOE3+ years SRE/DevOps experience, strong AWS and Terraform skills, Kubernetes and Helm experience, observability and incident response knowledge, bachelor's or equivalent, on-call participation, East Coast hours.
EarnIn: Provides immediate access to earned wages through a mobile app.
3+ YOE3+ years SRE or related experience; hands-on Python/Go coding; experience with observability, incident response, SLOs/SLIs, and distributed systems; strong communication and documentation skills.
Obsidian Security: Provides cybersecurity and threat detection for enterprise SaaS applications.
3+ YOE3+ years DevOps/SRE experience on GCP and/or AWS, Bachelor's in CS or related, proficiency with Kubernetes, Helm, GitLab CI/CD, ArgoCD, Prometheus, Grafana; programming in Golang or Python; strong communication and critical thinking.