44 site reliability manager jobs at 29 companies in Soquel, CA

1mo
Save
Mark Applied
Hide
Site Reliability Engineer
Santa Clara or St. Louis or Bangalore or London or Paris or Melbourne or Taipei or Tokyo
OnsiteFull Time
Netskope
NetskopeNASDAQ: NTSK: Cloud-native cybersecurity and data protection platform for enterprises.
3+ YOEBachelor's in CS/Engineering or equivalent; 3+ years building/managing complex systems (including 1-2 years SRE); experience with cloud services, microservices, availability/performance optimization, debugging, and strong communication.
Python, C, C++, Go, Rust, Docker, Kubernetes, AWS, GCP, KVM, OpenNebula, OpenStack, TCP/IP
3w
Save
Mark Applied
Hide
Senior Site Reliability Engineer
Sunnyvale, California, United States
$90k-$180k/yr OnsiteFull Time
Abbott
AbbottNYSE: ABT: Manufactures medical devices, diagnostics, and nutritional health products.
Ensure reliability, scalability, and performance of a medical-device remote monitoring platform; expertise in cloud (Azure), Kubernetes, observability, automation, and incident management; bachelor's in a technical discipline.
Python, Go, Bash, PowerShell, Microsoft Azure, Azure Kubernetes Service (AKS), Azure Monitor, Azure DevOps, Azure Policy, Kubernetes, Docker, Prometheus, Grafana, ELK/EFK, Datadog, Linux
4w
Save
Mark Applied
Hide
Site Reliability Engineer
Santa Clara, California, United States
$230k-$250k/yr OnsiteFull Time
Forward Networks
Forward Networks: Provides a digital twin platform for enterprise network management.
6+ YOE6+ years SRE/DevOps experience in SaaS/cloud, strong networking fundamentals, Kubernetes, observability (Prometheus/Grafana/Datadog/Splunk), Python/Bash automation, cloud and IaC (AWS/GCP/Azure, Terraform/Ansible), and incident response ownership.
Kubernetes, Prometheus, Grafana, Datadog, Splunk, Python, Bash, AWS, GCP, Azure, Terraform, Ansible
3w
Save
Mark Applied
Hide
Staff Site Reliability Engineer
Santa Clara, California, United States
$163k-$214k/yr OnsiteFull Time
IonQ
IonQNYSE: IONQ: Develops and sells trapped-ion quantum computers and cloud services.
7+ YOE7+ years production engineering experience; hands-on AWS/GCP reliability, observability and SLO ownership, incident command, resilience testing, and multi-team technical leadership.
AWS, GCP, Amazon Bedrock Agent Core
1w
Save
Mark Applied
Hide
Senior Site Reliability Engineer - Cloud
Santa Clara, California, United States
$168k-$265k/yr OnsiteFull Time
NVIDIA
NVIDIANASDAQ: NVDA: Designs graphics processing units and artificial intelligence hardware.
8+ YOEMS/BS or equivalent experience, 8+ years supporting live-site production, SRE on-call experience, strong Kubernetes and Python skills, Akamai/CDN and AWS experience, incident management and automation focus.
Akamai Edge Redirector Cloudlets, Akamai Forward Rewrite Cloudlets, Akamai Cloudlets Policy Manager, Akamai CDN, WAF, AWS, Kubernetes, Python
1w
Save
Mark Applied
Hide
Site Reliability Engineering (SRE) Manager, Apple Maps
Cupertino, California, United States
OnsiteFull Time
Apple
AppleNASDAQ: AAPL: Designs and sells consumer electronics, software, and online services.
Build, manage, and deliver highly available, automated infrastructure for Apple Maps at global scale; focus on reliability, scalability, and operational excellence.
3w
Save
Mark Applied
Hide
Senior Site Reliability Engineer
Sunnyvale or Sylmar
$90k-$180k/yr OnsiteFull Time
Abbott
AbbottNYSE: ABT: Provides medical devices, diagnostics, and science-based nutritional products.
Senior SRE with strong distributed systems, cloud (Azure), Kubernetes, observability, automation, incident management, and cross-functional communication skills for a medical device remote monitoring platform.
Python, Go, Bash, PowerShell, Microsoft Azure, Azure Kubernetes Service (AKS), Azure Monitor, Azure DevOps, Azure Policy, Kubernetes, Docker, Prometheus, Grafana, ELK, EFK, Datadog, Linux
1mo
Save
Mark Applied
Hide
Sr. Site Reliability Engineer
Palo Alto or Palo Alto or Washington
$165k-$230k/yr OnsiteFull Time
SpaceX
SpaceX: Designs and launches advanced rockets and satellite internet constellations.
5+ YOE5+ years experience with Kubernetes and Linux, proficiency in Bash/Python, experience with infrastructure automation and large-scale server management; Top Secret/SCI clearance required or obtainable.
Kubernetes, Linux, Bash, Python, Bazel, Makefiles, Terraform, Ansible, TCP/IP
2mo
Save
Mark Applied
Hide
Staff Site Reliability Engineer
Foster City, California, United States
$250k-$300k/yr HybridFull Time
Zoox
ZooxNASDAQ: AMZN: Developing autonomous robotaxis for urban ride-hailing services.
5+ YOE5+ years operating GitHub Enterprise at scale, monorepo management, CI/CD integration, infrastructure-as-code (Terraform/Pulumi), cloud platform experience, technical leadership and migration planning.
Git, GitHub Enterprise, GitHub Cloud, Buildkite, GitHub Actions, Jenkins, GitLab CI, Terraform, Pulumi, Bazel, Buck, Reviewable, Gerrit
2mo
Save
Mark Applied
Hide
Staff Site Reliability Engineer
Mountain View, California, United States
$252k-$308k/yr HybridFull Time
EarnIn
EarnIn: Provides immediate access to earned wages through a mobile app.
7+ YOE7+ years in SRE or related field; experience applying AI/LLMs to operations; strong SLO/SLI and incident management; software engineering in Python or Go; observability and IaC proficiency; AI-assisted development tools; fintech/regulated environment experience.
Datadog, CloudWatch, OpenTelemetry, Terraform, Kubernetes, AWS, Python, Go, Cursor, Claude Code, Copilot
2mo
Save
Mark Applied
Hide
Sr. Site Reliability Engineer
Sunnyvale, California, United States
$170k-$196k/yr OnsiteFull Time
Illumio
Illumio: Provides zero-trust segmentation software to contain cyberattacks.
5+ YOE5+ years SRE experience with AWS and/or Azure, scripting in PowerShell/Python/Go, CI/CD experience (Azure DevOps, Jenkins, GitLab CI/CD), containerization knowledge (Docker, Kubernetes), bachelor’s degree or equivalent, on-call and incident management experience.
AWS, Azure, PowerShell, Python, Go, Azure DevOps, Jenkins, GitLab CI/CD, Docker, Kubernetes
4w
Save
Mark Applied
Hide
Senior Lead Site Reliability Engineer
Palo Alto, California, United States
$171k-$260k/yr OnsiteFull Time
JPMorgan Chase
JPMorgan ChaseNYSE: JPM: Global financial services firm providing banking and investment solutions.
5+ YOE5+ years applied SRE experience, formal SRE training/certification, expertise in observability, distributed systems, cloud-native and AI-assisted reliability workflows.
Grafana, Dynatrace, Prometheus, Datadog, Splunk, Java, Go (Golang), Python, Terraform, LangChain, LangGraph, AutoGen, CrewAI, GitHub Copilot, Claude, Fluentd, Logstash, Vector, Kafka, RabbitMQ, SQS, Neo4j, TigerGraph, Pinecone, Weaviate, Chroma, Docker, Kubernetes, GitOps, MCP (Model Context Protocol), TensorFlow, PyTorch, scikit-learn, Hadoop, Spark, Flink, MongoDB, Cassandra, DynamoDB, InfluxDB, TimescaleDB, Chaos Monkey, Gremlin, LitmusChaos
1mo
Save
Mark Applied
Hide
Site Reliability Engineer, Compute Platform
San Jose, California, United States
OnsiteFull Time
ByteDance
ByteDance: Developing AI-driven content platforms and mobile applications.
Experience with Linux, networking, databases, Kubernetes, ClickHouse/Hadoop/Doris/Spark/Presto, scripting or programming (Python, Shell, Java, Go), incident management, and capacity planning.
ClickHouse, Spark, Presto, Doris, Hadoop, Kubernetes, Linux, Python, Shell, Java, Go
2mo
Save
Mark Applied
Hide
Principal Site Reliability Engineer
Santa Clara, California, United States
$152k-$245k/yr OnsiteFull Time
Palo Alto Networks
Palo Alto NetworksNASDAQ: PANW: Provides enterprise-grade network, cloud, and endpoint security software.
BS or MS in CS or related field; expertise in configuration management (Ansible, Terraform, Kubernetes); Python and/or Go; Kubernetes with autoscaling; production engineering/DevOps/SRE experience; public cloud (GCP/AWS); Linux networking; CI/CD with GitLab/GitHub; distributed systems; strong communication; ownership and monitoring as code.
Kubernetes, Docker, GCP, AWS, Ansible, Terraform, Vault, GitLab, Spinnaker, Pub/Sub, Bigtable, Memorystore, BigQuery, RabbitMQ, Kafka, MySQL, Python, Go, Shell scripting, Golang
2mo
Save
Mark Applied
Hide
Site Reliability Engineer – USDS (Multiple Positions)
San Jose, California, United States
$188k-$259k/yr OnsiteFull Time
TikTok USDS Joint Venture
TikTok USDS Joint Venture: Operates and secures TikTok services for U.S. users.
1+ YOEDegree in CS/Engineering/IT/Math plus related experience (Master's+1yr or Bachelor's+3yrs); experience monitoring, troubleshooting, SLA management, runbooks, incident response and postmortems.
2mo
Save
Mark Applied
Hide
Senior Software Engineer, Site Reliability Engineering
San Francisco or San Jose or New York City or Seattle or Austin or Washington or California or Massachusetts or New Jersey or Washington or United States
$179k-$273k/yr RemoteFull Time
Thumbtack
Thumbtack: Online marketplace connecting homeowners with local service professionals.
5+ YOE5+ years managing infrastructure and systems; extensive AWS and Linux fluency; proficiency in Python, Go, PHP, and JavaScript; experience with distributed systems, observability, and on-call rotations; strong communication and troubleshooting skills.
AWS, Linux, Python, Go, PHP, JavaScript, DNS, TLS, HTTP/S, TCP/IP
5d
Save
Mark Applied
Hide
Sr. Site Reliability Engineer - Core Platform & Embedded Reliability (Hybrid)
New York City or Austin or Sunnyvale or Redmond
$140k-$215k/yr HybridFull Time
CrowdStrike
CrowdStrikeNASDAQ: CRWD: Provides cloud-native endpoint protection and cybersecurity services.
10+ YOE10+ years building distributed systems, 5+ years developing SaaS microservices, expert programming skills, distributed-systems expertise, architectural leadership, and a Computer Science degree or equivalent experience.
Go, Java, Scala, Kotlin, Python, Node.js, Kubernetes, AWS, Cassandra, Kafka, Elasticsearch, OpenSearch, Google Cloud Platform (GCP), Oracle Cloud Infrastructure (OCI), GitHub, Stack Overflow
2w
Save
Mark Applied
Hide
Staff Site Reliability Engineer-Production Operations
Palo Alto, California, United States
$186k-$233k/yr HybridFull Time
Rivian and Volkswagen Group Technologies
Rivian and Volkswagen Group Technologies: A joint venture creating software-defined vehicle technology and connected services for electric vehicles.
Senior SRE with incident command experience, strong systems engineering for distributed systems, hands-on coding (Python or Go), observability expertise (Datadog or comparable), and experience with blameless post-incident practices.
Python, Go, Datadog, LLM
1mo
Save
Mark Applied
Hide
Contract Lead, Site Reliability Engineering — AI Accelerator Infrastructure
Santa Clara, California, United States
$195k-$285k/yr HybridContract, Full Time
d-Matrix: Develops high-performance semiconductor chips for generative AI inference.
15+ YOE5+ MgmtBachelor's in CS/EE,15+ years SRE/infrastructure engineering,5+ years leading SRE teams,deep Linux,Terraform,Ansible,Kubernetes,Prometheus/Grafana/Datadog,Python or Go,cloud (AWS/Azure/GCP).
Prometheus, Grafana, Datadog, Terraform, Ansible, Kubernetes, Python, Go, AWS, Azure, GCP, Slurm, LSF, InfiniBand, RoCE, NVLink
1mo
Save
Mark Applied
Hide
Tech Lead Site Reliability Engineer, TikTok Generalized Arch USTO
San Jose, California, United States
$245k-$450k/yr OnsiteFull Time
TikTok
TikTok: Global short-form video hosting and social media platform.
5+ YOEBachelor's in CS or related, strong CS foundation, Linux and storage/network knowledge, proficiency in Python/Go/Java/PHP/C/C++, strong problem solving and communication; 5+ years SRE/cloud experience preferred.
Linux, Python, Go, Java, PHP, C, C++

Explore Jobs

Expand Your Job Search