41 site reliablity engineer jobs at 23 companies in Thousand Oaks, CA
3mo
Save
Mark Applied
Hide
3mo
Site Reliability Engineer
Europe or Canada or Bellevue or Los Angeles
RemoteFull Time
ShopifyNYSE: SHOP: A global commerce platform for building and managing online stores.
Experienced SRE/engineer with on-call experience, ability to build resilient production tooling, improve observability, respond to alerts, and collaborate across engineering teams.
Green DotNYSE: GDOT: Provides mobile banking and payment solutions to consumers and businesses.
7+ YOE7+ years in release/reliability engineering, cloud platform experience (AWS/Azure/GCP), automated deployment and observability proficiency, scripting with PowerShell/Bash/Python, excellent troubleshooting and communication skills.
SpaceX: Designs and launches advanced rockets and satellite internet constellations.
1+ YOE1+ years hands-on experience with client/server hardware, networking, Linux/Windows, scripting and automation; bachelor's in CS/engineering/math or 2+ years software experience in lieu; HPC and systems engineering experience preferred.
K2 Space: Develops high-power satellite platforms for heavy-lift launch vehicles.
5+ YOE5+ years SRE/DevOps experience or BS in CS/IT/STEM, deep cloud (AWS/GCP/Azure), IaC, Kubernetes, Linux, programming (Go/Python), strong security and reliability experience.
Picogrid: Develops hardware and software for autonomous defense systems.
3+ YOE3+ years SRE experience, deep Kubernetes and Terraform/OpenTofu skills, AWS proficiency, observability (Grafana, Prometheus, Loki, OpenTelemetry), incident response, HA databases, and IoT/edge fleet experience.
The Walt Disney CompanyNYSE: DIS: Produces movies, operates theme parks, and provides streaming services.
10+ YOE10+ years experience; expertise in multi-cloud (AWS, Azure, GCP), observability, CI/CD, infrastructure as code (Terraform/CloudFormation), Linux systems administration; bachelor's degree or equivalent; strong leadership and communication.
MicrosoftNASDAQ: MSFT: Develops computer software, consumer electronics, and video games.
1+ YOE1–3 years SRE/DevOps or cloud experience; familiarity with Linux, HTTP, DNS, containers, Kubernetes, Git, and scripting (Bash/Python); exposure to monitoring, logs, metrics, and incident management; strong troubleshooting and communication skills.
AbbottNYSE: ABT: Provides medical devices, diagnostics, and science-based nutritional products.
Senior SRE with strong distributed systems, cloud (Azure), Kubernetes, observability, automation, incident management, and cross-functional communication skills for a medical device remote monitoring platform.
Python, Go, Bash, PowerShell, Microsoft Azure, Azure Kubernetes Service (AKS), Azure Monitor, Azure DevOps, Azure Policy, Kubernetes, Docker, Prometheus, Grafana, ELK, EFK, Datadog, Linux
AXS: Provides digital ticketing and marketing solutions for live events.
4+ YOE4-6 years in site reliability or DevOps; BA/BS preferred but not required; cloud operations, infrastructure as code, containers/orchestration, CI/CD; programming/scripting ability to automate tasks.
Cloud, Containers, Orchestration, Infrastructure as Code, CI/CD, Python, Bash, Go
Fox CorporationNasdaq: FOXA: Broadcasts news, sports, and entertainment via television and streaming.
5+ YOEExpertise with EKS/Kubernetes/Istio and AWS, 5+ years SRE/DevOps experience, strong AI/ML knowledge, proficiency with Git and Terraform, programming in Golang or Python, troubleshooting APIs and performing root-cause analysis.
Hadrian: Building autonomous factories for aerospace and defense manufacturing.
Own the reliability of production robotics systems; Kubernetes, telemetry, and coding in TypeScript, Python, Golang, or C++. Design SLOs/SLIs and implement observability.
HiveWatch: Cloud-based physical security platform for enterprise operations.
5+ YOE5+ years software engineering and 3+ years SRE, DevOps, or production operations experience; AWS, Docker, Kubernetes, IaC, SQL optimization, observability, distributed systems, and debugging skills required.
Site Reliability Engineer Intern (Global SRE) - 2027 Summer
San Jose or Los Angeles or New York City or London or Dublin or Paris or Berlin or Dubai or Jakarta or Seoul or Tokyo
OnsiteInternship, Full Time
TikTok: Global short-form video hosting and social media platform.
Currently pursuing a bachelor's degree in computer science or related field; Unix/Linux, IP networking, and Python, Go, C, C++, or Java programming experience required.
Unix/Linux, IP networking, Python, Go, C, C++, Java
The Aerospace Corporation: Provides technical expertise and research for national space programs.
8+ YOE8+ years in distributed software; 5+ years with Kubernetes in enterprise; 2+ years managing Kubernetes; Linux admin; automation; TS/SCI clearance preferred
Lightspark: Builds real-time global payment infrastructure on the Lightning Network
5+ YOERequires 5+ years of production, site reliability, or DevOps engineering experience, tested-code expertise, cross-functional collaboration, and cloud infrastructure familiarity. A CS degree is ideal but not required.