179 site reliability engineering jobs at 95 companies in San Jose, CA

1mo
Save
Mark Applied
Hide
Manager, Site Reliability Engineering
San Francisco, California, United States
$204k-$306k/yr HybridFull Time
Okta
OktaNASDAQ: OKTA: Provide secure identity management and authentication for enterprises.
3+ Mgmt3+ years technical leadership experience; experience with cloud-native architectures, Kubernetes, Terraform, CI/CD, observability platforms; strong software development and automation background; US Person status required.
Amazon Web Services (AWS), Kubernetes, Terraform, Grafana, Splunk, APM, CI/CD
3mo
Save
Mark Applied
Hide
Site Reliability Engineering
Foster City, California, United States
$140k-$230k/yr HybridFull Time
Zoox
ZooxNASDAQ: AMZN: Developing autonomous robotaxis for urban ride-hailing services.
5+ YOE5+ years SRE/Distributed systems; cloud platforms (AWS, GCP, or Azure); IaC (Terraform, Ansible, Salt, CloudFormation); Kubernetes; Python/Go/C/C++/Java.
AWS, GCP, Azure, Terraform, Ansible, Salt, CloudFormation, Kubernetes, Python, Go, C/C++, Java
2mo
Save
Mark Applied
Hide
Site Reliability Engineer
San Francisco or South San Francisco
$150k/yr OnsiteFull Time
VantageScore
VantageScore: Provides credit scoring and data analytics solutions.
5+ YOEExperienced Site Reliability Engineer with a DevSecOps focus; patch management, vulnerability remediation; AWS and CI/CD, security tooling.
AWS, EC2, ECS, Lambda, EKS, S3, RDS, IAM, VPC, CloudTrail, Config, GuardDuty, GitHub Actions, CodePipeline, Terraform, CloudFormation, AWS CDK, Kubernetes, Snyk, Wiz, Prisma Cloud, Kong, HashiCorp Vault, Secrets Manager, CloudWatch, Datadog, Grafana
2mo
Save
Mark Applied
Hide
Lead Site Reliability Engineering - Network
Palo Alto or Columbus
$152k-$215k/yr OnsiteFull Time
JPMorgan Chase
JPMorgan ChaseNYSE: JPM: Global financial services firm providing banking and investment solutions.
5+ YOE10+ MgmtFormal network engineering training, 5+ years applied experience, 10+ years leading technologists, advanced network reliability skills, SD-WAN and cloud (AWS, Azure) proficiency, major network vendor experience, observability tooling and incident leadership.
SD-WAN, AWS, Azure, Palo Alto, Juniper, F5, Broadcom, Arista, Cisco, Grafana, SevOne, Prometheus, Kibana, ThousandEyes, Splunk, Jenkins, GitLab, Terraform, eBPF, TCP/IP, HTTPS, BGP
5d
Save
Mark Applied
Hide
Site Reliability Engineer (SRE)
Santa Clara, California, United States
$50-$60/hr RemoteContract
ServiceNow
ServiceNowNYSE: NOW: Enterprise cloud platform for digital workflow automation.
3+ YOEBachelor's degree in computer science or related field; 3+ years in site reliability engineering; 2+ years with AWS and cloud automation; Kubernetes, Linux, Terraform, networking, GitOps, monitoring, and customer support experience.
AWS, Kubernetes, Helm, Linux, Terraform, GitOps, Prometheus, Grafana, Bazel, CueLang, Version Control, Okta, Snowflake, Google
1w
Save
Mark Applied
Hide
Site Reliability Engineer (SRE)
San Francisco, California, United States
$350k-$475k/yr OnsiteFull Time
Thinking Machines
Thinking Machines: Building AI systems to extend human will and judgment.
Experience in distributed systems/cloud/site reliability, software automation for reliability, incident response and postmortems, strong communication and coordination skills.
Tinker, Kubernetes, LoRA, CI/CD
2mo
Save
Mark Applied
Hide
Senior Site Reliability Engineer
New York City or Austin or Berlin or Bucharest or Chicago or Dubai or Jakarta or London or Paris or San Francisco or São Paulo or Singapore or Seoul or Sydney or Tokyo
HybridFull Time
Braze
BrazeNASDAQ: BRZE: Platform for personalized customer engagement and cross-channel messaging.
3+ YOE3+ years as a Software/DevOps/Site Reliability Engineer, strong Linux/Unix shell skills, programming experience in Ruby and/or Go, experience with Docker, Kubernetes, Terraform/Chef, and data stores like MongoDB, Redis, Kafka, or Postgres.
Ruby on Rails, Ruby, Go, Linux, Unix Shell, Docker, Kubernetes, Terraform, Chef, MongoDB, Redis, Kafka, Postgres, PagerDuty
3w
Save
Mark Applied
Hide
Site Reliability Engineer
San Francisco or Alpharetta or Arlington or Augusta or Ashburn or Allentown or Appleton or Atlanta or Annapolis Junction or Ann Arbor or Herndon or Allen
$165k-$241k/yr RemoteFull Time
Cisco
CiscoNASDAQ: CSCO: Develops and sells networking hardware and cybersecurity software.
7+ YOE7+ years SRE or related experience; BS/MS/PhD with corresponding years; U.S. Person required for FedRAMP/IL-5 work; on-call participation; strong coding, automation, reliability, and security skills.
1mo
Save
Mark Applied
Hide
Site Reliability Engineer
San Francisco, California, United States
HybridFull Time
Runloop
Runloop: Provides infrastructure and secure sandboxes for AI agents.
5+ YOE5+ years software engineering experience with 3+ years in SRE/DevOps, strong Python or Go skills, containerization, cloud infra, monitoring, networking, Linux administration, on‑call and incident management.
AWS, GCP, Azure, Grafana, Prometheus, Datadog, Python, Go, Docker, Kubernetes, Terraform, Pulumi, Sentry, RUM, CI/CD
1mo
Save
Mark Applied
Hide
Site Reliability Engineer
Santa Clara or St. Louis or Bangalore or London or Paris or Melbourne or Taipei or Tokyo
OnsiteFull Time
Netskope
NetskopeNASDAQ: NTSK: Cloud-native cybersecurity and data protection platform for enterprises.
3+ YOEBachelor's in CS/Engineering or equivalent; 3+ years building/managing complex systems (including 1-2 years SRE); experience with cloud services, microservices, availability/performance optimization, debugging, and strong communication.
Python, C, C++, Go, Rust, Docker, Kubernetes, AWS, GCP, KVM, OpenNebula, OpenStack, TCP/IP
1mo
Save
Mark Applied
Hide
Senior Site Reliability Engineer
Oakland, California, United States
$175k-$210k/yr HybridFull Time
Fivetran
Fivetran: Automates data movement into cloud data warehouses.
5+ YOE5+ years SaaS experience; managed Kubernetes, cloud platforms (AWS/GCP/Azure), Terraform/Ansible/ArgoCD; Python/Shell scripting, Linux admin, PostgreSQL; incident response and reliability engineering experience.
Kubernetes, EKS, AKS, GKE, PostgreSQL, ArgoCD, Terraform, Ansible, Python, Shell, Go, Java, AWS, GCP, Azure, Grafana, Buildkite, Temporal, Pulumi, Linux, VPN, PrivateLink, Private Service Connect (GCP)
1mo
Save
Mark Applied
Hide
Director, Engineering Operations and Site Reliability Engineering - Datacenter Server Systems
Santa Clara, California, United States
$292k-$443k/yr OnsiteFull Time
NVIDIA
NVIDIANASDAQ: NVDA: Designs GPU-accelerated computing and artificial intelligence hardware.
12+ YOE7+ MgmtBS/MS in CS/EE/CE or equivalent; 12+ years in infrastructure/systems/reliability/datacenter ops, including 7+ years people management; strong Linux, server, networking, monitoring, incident management, automation skills; strong communication.
Linux, NVLink, InfiniBand, Spectrum-X
2w
Save
Mark Applied
Hide
Site Reliability Engineering (SRE) Manager, Apple Maps
Cupertino, California, United States
OnsiteFull Time
Apple
AppleNASDAQ: AAPL: Designs and sells consumer electronics, software, and online services.
Build, manage, and deliver highly available, automated infrastructure for Apple Maps at global scale; focus on reliability, scalability, and operational excellence.
1mo
Save
Mark Applied
Hide
Senior Software Engineer, Site Reliability Engineering
New York or San Ramon or Reno
$153k-$210k/yr HybridFull Time
Ridgeline
Ridgeline: Cloud-native platform for investment management operations.
3+ YOE3–6 years SRE/DevOps experience, 2+ years on AWS, proficiency with Terraform, observability, CI/CD, Python/Go/Bash, incident response, and strong communication and troubleshooting skills.
Claude Code, Cursor, Terraform, AWS, EC2, ECS, EKS, RDS, S3, IAM, CloudWatch, GitHub Actions, CircleCI, Buildkite, Python, Go, Bash, Kubernetes, Helm, Kotlin, Node.js, TypeScript
2mo
Save
Mark Applied
Hide
Senior Software Engineer, Site Reliability Engineering
San Francisco or San Jose or New York City or Seattle or Austin or Washington or California or Massachusetts or New Jersey or Washington or United States
$179k-$273k/yr RemoteFull Time
Thumbtack
Thumbtack: Online marketplace connecting homeowners with local service professionals.
5+ YOE5+ years managing infrastructure and systems; extensive AWS and Linux fluency; proficiency in Python, Go, PHP, and JavaScript; experience with distributed systems, observability, and on-call rotations; strong communication and troubleshooting skills.
AWS, Linux, Python, Go, PHP, JavaScript, DNS, TLS, HTTP/S, TCP/IP
2mo
Save
Mark Applied
Hide
Founding Engineer - Site Reliability
San Francisco or United States
$185k-$285k/yr RemoteFull Time
uRun
uRun: Infrastructure cloud for interactive, stateful AI inference.
7+ YOE7+ years in site reliability or infrastructure engineering; strong SLOs, incident response, and observability; Kubernetes and cloud (AWS); software engineering fundamentals; first SRE at a company.
Kubernetes, AWS, Prometheus, Grafana, Datadog, Automation, VPC, GPU compute
1mo
Save
Mark Applied
Hide
Site Reliability Engineer
Santa Clara, California, United States
$230k-$250k/yr OnsiteFull Time
Forward Networks
Forward Networks: Provides a digital twin platform for enterprise network management.
6+ YOE6+ years SRE/DevOps experience in SaaS/cloud, strong networking fundamentals, Kubernetes, observability (Prometheus/Grafana/Datadog/Splunk), Python/Bash automation, cloud and IaC (AWS/GCP/Azure, Terraform/Ansible), and incident response ownership.
Kubernetes, Prometheus, Grafana, Datadog, Splunk, Python, Bash, AWS, GCP, Azure, Terraform, Ansible
2w
Save
Mark Applied
Hide
Site Reliability Engineer
San Mateo or Arizona or California or Colorado or Florida or Georgia or Illinois or Nevada or North Carolina or Oregon or Texas or Utah or Washington
$140k-$150k/yr RemoteFull Time
VyncaCare
VyncaCare: Offers palliative care services and advance care planning technology.
3+ YOE3+ years SRE/DevOps experience, strong AWS and Terraform skills, Kubernetes and Helm experience, observability and incident response knowledge, bachelor's or equivalent, on-call participation, East Coast hours.
AWS, Terraform, Kubernetes, Helm, Prometheus, Grafana, Datadog, CloudWatch, SigNoz, OpenTelemetry, ArgoCD, Flux, PostgreSQL, MySQL, Redshift, ClickHouse, AWS Secrets Manager, HashiCorp Vault, Snowflake, Python, Go, Linux
1mo
Save
Mark Applied
Hide
Site Reliability Engineer
Mountain View, California, United States
$189k-$232k/yr HybridFull Time
EarnIn
EarnIn: Provides immediate access to earned wages through a mobile app.
3+ YOE3+ years SRE or related experience; hands-on Python/Go coding; experience with observability, incident response, SLOs/SLIs, and distributed systems; strong communication and documentation skills.
Python, Go, Datadog, CloudWatch, logs, metrics, traces, APM, GitHub Copilot, Cursor, ChatGPT, Claude
1mo
Save
Mark Applied
Hide
Site Reliability Engineer
Palo Alto or Newport Beach
$165k-$190k/yr OnsiteFull Time
Obsidian Security
Obsidian Security: Provides cybersecurity and threat detection for enterprise SaaS applications.
3+ YOE3+ years DevOps/SRE experience on GCP and/or AWS, Bachelor's in CS or related, proficiency with Kubernetes, Helm, GitLab CI/CD, ArgoCD, Prometheus, Grafana; programming in Golang or Python; strong communication and critical thinking.
Kubernetes, Helm, GitLab CI/CD, ArgoCD, Prometheus, Grafana, Golang, Python, Kafka, Elasticsearch, PostgreSQL, ScyllaDB, Databricks, Dagster, Sentry, Kong, AWS, GCP