44 site reliability manager jobs at 29 companies in Soquel, CA
1mo
Save
Mark Applied
Hide
1mo
Site Reliability Engineer
Santa Clara or St. Louis or Bangalore or London or Paris or Melbourne or Taipei or Tokyo
OnsiteFull Time
NetskopeNASDAQ: NTSK: Cloud-native cybersecurity and data protection platform for enterprises.
3+ YOEBachelor's in CS/Engineering or equivalent; 3+ years building/managing complex systems (including 1-2 years SRE); experience with cloud services, microservices, availability/performance optimization, debugging, and strong communication.
AbbottNYSE: ABT: Manufactures medical devices, diagnostics, and nutritional health products.
Ensure reliability, scalability, and performance of a medical-device remote monitoring platform; expertise in cloud (Azure), Kubernetes, observability, automation, and incident management; bachelor's in a technical discipline.
Python, Go, Bash, PowerShell, Microsoft Azure, Azure Kubernetes Service (AKS), Azure Monitor, Azure DevOps, Azure Policy, Kubernetes, Docker, Prometheus, Grafana, ELK/EFK, Datadog, Linux
Site Reliability Engineering (SRE) Manager, Apple Maps
Cupertino, California, United States
OnsiteFull Time
AppleNASDAQ: AAPL: Designs and sells consumer electronics, software, and online services.
Build, manage, and deliver highly available, automated infrastructure for Apple Maps at global scale; focus on reliability, scalability, and operational excellence.
AbbottNYSE: ABT: Provides medical devices, diagnostics, and science-based nutritional products.
Senior SRE with strong distributed systems, cloud (Azure), Kubernetes, observability, automation, incident management, and cross-functional communication skills for a medical device remote monitoring platform.
Python, Go, Bash, PowerShell, Microsoft Azure, Azure Kubernetes Service (AKS), Azure Monitor, Azure DevOps, Azure Policy, Kubernetes, Docker, Prometheus, Grafana, ELK, EFK, Datadog, Linux
SpaceX: Designs and launches advanced rockets and satellite internet constellations.
5+ YOE5+ years experience with Kubernetes and Linux, proficiency in Bash/Python, experience with infrastructure automation and large-scale server management; Top Secret/SCI clearance required or obtainable.
EarnIn: Provides immediate access to earned wages through a mobile app.
7+ YOE7+ years in SRE or related field; experience applying AI/LLMs to operations; strong SLO/SLI and incident management; software engineering in Python or Go; observability and IaC proficiency; AI-assisted development tools; fintech/regulated environment experience.
BS or MS in CS or related field; expertise in configuration management (Ansible, Terraform, Kubernetes); Python and/or Go; Kubernetes with autoscaling; production engineering/DevOps/SRE experience; public cloud (GCP/AWS); Linux networking; CI/CD with GitLab/GitHub; distributed systems; strong communication; ownership and monitoring as code.
Site Reliability Engineer – USDS (Multiple Positions)
San Jose, California, United States
$188k-$259k/yrOnsiteFull Time
TikTok USDS Joint Venture: Operates and secures TikTok services for U.S. users.
1+ YOEDegree in CS/Engineering/IT/Math plus related experience (Master's+1yr or Bachelor's+3yrs); experience monitoring, troubleshooting, SLA management, runbooks, incident response and postmortems.
Senior Software Engineer, Site Reliability Engineering
San Francisco or San Jose or New York City or Seattle or Austin or Washington or California or Massachusetts or New Jersey or Washington or United States
$179k-$273k/yrRemoteFull Time
Thumbtack: Online marketplace connecting homeowners with local service professionals.
5+ YOE5+ years managing infrastructure and systems; extensive AWS and Linux fluency; proficiency in Python, Go, PHP, and JavaScript; experience with distributed systems, observability, and on-call rotations; strong communication and troubleshooting skills.
CrowdStrikeNASDAQ: CRWD: Provides cloud-native endpoint protection and cybersecurity services.
10+ YOE10+ years building distributed systems, 5+ years developing SaaS microservices, expert programming skills, distributed-systems expertise, architectural leadership, and a Computer Science degree or equivalent experience.
Staff Site Reliability Engineer-Production Operations
Palo Alto, California, United States
$186k-$233k/yrHybridFull Time
Rivian and Volkswagen Group Technologies: A joint venture creating software-defined vehicle technology and connected services for electric vehicles.
Senior SRE with incident command experience, strong systems engineering for distributed systems, hands-on coding (Python or Go), observability expertise (Datadog or comparable), and experience with blameless post-incident practices.
Contract Lead, Site Reliability Engineering — AI Accelerator Infrastructure
Santa Clara, California, United States
$195k-$285k/yrHybridContract, Full Time
d-Matrix: Develops high-performance semiconductor chips for generative AI inference.
15+ YOE5+ MgmtBachelor's in CS/EE,15+ years SRE/infrastructure engineering,5+ years leading SRE teams,deep Linux,Terraform,Ansible,Kubernetes,Prometheus/Grafana/Datadog,Python or Go,cloud (AWS/Azure/GCP).
Tech Lead Site Reliability Engineer, TikTok Generalized Arch USTO
San Jose, California, United States
$245k-$450k/yrOnsiteFull Time
TikTok: Global short-form video hosting and social media platform.
5+ YOEBachelor's in CS or related, strong CS foundation, Linux and storage/network knowledge, proficiency in Python/Go/Java/PHP/C/C++, strong problem solving and communication; 5+ years SRE/cloud experience preferred.