Blizzard Entertainment
Posted 2mo ago

Senior Site Reliability Engineer, Data & Analytics

Blizzard Entertainment
Irvine or Albany or United States
$101k-$187k/yrHybridFull Time
Responsibilities
  • resolving incidents
  • building automation
  • operating Kubernetes
Requirements
  • Production SRE experience operating distributed data/ML workloads
  • Strong Linux
  • Containers/Kubernetes
  • Terraform
  • CI/CD (Jenkins
  • GitHub Actions
  • ArgoCD)
  • Automation with Python/Go/shell
  • Observability and SRE practices
  • On-call participation
Technical tools mentioned
LinuxKubernetesTerraformJenkinsGitHub ActionsArgoCDPythonGoshellPrometheusGrafanaKafkaPub/SubGCPAWS

Job description

Team Name:

Battle.net & Online Products

Job Title:

Senior Site Reliability Engineer, Data & Analytics

Requisition ID:

R027436

Job Description:

This Senior Site Reliability Engineer role is on our Data & Analytics team, partnering with data, analytics, ML, and platform engineering to improve the reliability, scalability, and performance of large-scale data platforms, analytics pipelines, ML training pipelines, and inference services. 

In addition to core SRE responsibilities, this role will build operational and automation tooling that reduces toil, speeds up issue resolution, and improves engineering velocity. This includes contributing to internal platform services such as shared tooling, data integrations, and access-control patterns used across Blizzard. 

The ideal candidate is a production-minded SRE or platform engineer who is comfortable operating critical systems, writing software, and building tools that improve engineering efficiency without compromising reliability. 

This role is open to candidates based in Irvine, CA or Albany, NY (hybrid or on-site), as well as fully remote candidates. 

Responsibilities 

  • Participate in an on-call rotation and drive incidents to resolution 

  • Lead blameless postmortems and identify systemic reliability improvements 

  • Partner with data, ML, and platform teams to improve batch, streaming, training, and inference workloads 

  • Support ML training pipelines and inference services, including GPU workloads 

  • Help define how data and ML services run on Kubernetes 

  • Design and build automation and operational tooling (e.g., workflows, diagnostic tooling, runbooks) to reduce on-call burden 

  • Build and evolve centralized platform services, including shared tooling, data integrations, and access controls 

  • Diagnose and resolve reliability, performance, and cost issues across distributed systems 

  • Champion automation, documentation, and practices that reduce toil 

  • Maintain infrastructure using Terraform and infrastructure-as-code principles 

  • Improve CI/CD and GitOps workflows (Jenkins, GitHub Actions, ArgoCD) 

  • Operate and improve containerized services on Kubernetes 

  • Define and measure reliability using SLIs, SLOs, and error budgets 

  • Run load tests, capacity modeling, and production validation 

  • Build internal tools and paved paths that help teams operate safely and efficiently 

 

Minimum Requirements 

  • Experience operating reliable, distributed systems in SRE, platform, or similar roles 

  • Experience with data, analytics, ML, or large-scale distributed workloads 

  • Strong knowledge of Linux, containers, Kubernetes, and cloud infrastructure 

  • Experience building automation or internal tools (Python, Go, shell, etc.) 

  • Experience with infrastructure-as-code (e.g., Terraform) 

  • Experience with CI/CD or GitOps systems (e.g., Jenkins, GitHub Actions, ArgoCD) 

  • Familiarity with observability (metrics, logs, traces, alerting, incident response) 

  • Solid understanding of SRE concepts (SLIs, SLOs, error budgets, postmortems) 

  • Experience using modern development and automation practices to improve reliability and efficiency 

  • Experience building internal tooling, automation, or developer productivity systems 

  • Strong communication skills with technical and cross-functional partners 

 

Bonus Points 

  • Experience with data and ML systems (training pipelines, model serving, GPU workloads) 

  • Experience with distributed systems and messaging (Kafka, Pub/Sub) 

  • Experience working in Kubernetes-based environments 

  • Familiarity with observability tools (Prometheus, Grafana) 

  • Experience operating systems in cloud environments (GCP, AWS) 

 

Your Platform 

Best known for iconic video game universes including Warcraft®, Overwatch®, Diablo®, and StarCraft®, Blizzard Entertainment, Inc. (www.blizzard.com), a division of Activision Blizzard, which was acquired by Microsoft (NASDAQ: MSFT), is a premier developer and publisher of entertainment experiences. Blizzard Entertainment has created some of the industry’s most critically acclaimed and genre-defining games over the last 30 years, with a track record that includes multiple Game of the Year awards. Blizzard Entertainment engages tens of millions of players around the world with titles available on PC via Battle.net®, Xbox, PlayStation, Nintendo Switch, iOS, and Android.  

   

Our World 

Activision Blizzard, Inc., is one of the world's largest and most successful interactive entertainment companies and is at the intersection of media, technology and entertainment. We are home to some of the most beloved entertainment franchises including Call of Duty®, World of Warcraft®, Overwatch®, Diablo®, Candy Crush and Bubble Witch™. Our combined entertainment network delights hundreds of millions of monthly active users in 196 countries, making us the largest gaming network on the planet!  

Our ability to build immersive and innovative worlds is only enhanced by diverse teams working in an inclusive environment. We aspire to have a culture where everyone can thrivein order toconnect and engage the world through epic entertainment. We provide a suite of benefits that promote physical,emotionaland financial well-being for Every World -wevegot our employees covered!

 

The video game industry and therefore our business is fast-paced and will continue to evolve. As such, the duties and responsibilities of this role may be changed as directed by the Company at any time to promote and support our business and relationships with industry partners

We love hearing from anyone who is enthusiastic about changing the games industry. Not sure you meet all the qualifications? Let us decide! Research shows that women and members of other under-represented groups tend to not apply to jobs when they think they may not meet every qualification, when, in fact, they often do! We are committed to creating a diverse and inclusive environment and strongly encourage you to apply. 

We are committed to working with and providing reasonable assistance to individuals with physical and mental disabilities. If you are a disabled individual requiring an accommodation to apply for an open position, please email your request to [email protected]. General employment questions cannot be accepted or processed here. Thank you for your interest.

We are an equal opportunity employer and value diversity at our company. We do not discriminate on the basis of race, religion, color, national origin, gender, sexual orientation, gender identity, age, marital status, veteran status, or disability status, among other characteristics. 

Rewards

We provide a suite of benefits that promote physical, emotional and financial well-being for ‘Every World’ - we’ve got our employees covered!  Subject to eligibility requirements, the Company offers comprehensive benefits including:

  • Medical, dental, vision, health savings account or health reimbursement account, healthcare spending accounts, dependent care spending accounts, life and AD&D insurance, disability insurance;
  • 401(k) with Company match, tuition reimbursement, charitable donation matching;
  • Paid holidays and vacation, paid sick time, floating holidays, compassion and bereavement leaves, parental leave;
  • Mental health & wellbeing programs, fitness programs, free and discounted games, and a variety of other voluntary benefit programs like supplemental life & disability, legal service, ID protection, rental insurance, and others;
  • If the Company requires that you move geographic locations for the job, then you may also be eligible for relocation assistance.

Eligibility to participate in these benefits may vary for part time and temporary full-time employees and interns with the Company.  You can learn more by visiting https://www.benefitsforeveryworld.com/.

In the U.S., the standard base pay range for this role is $101,000.00 - $186,754.00 Annual. These values reflect the expected base pay range of new hires across all U.S. locations. Ultimately, your specific range and offer will be based on several factors, including relevant experience, performance, and work location. Your Talent Professional can share this role’s range details for your local geography during the hiring process. In addition to a competitive base pay, employees in this role may be eligible for incentive compensation. Incentive compensation is not guaranteed. While we strive to provide competitive offers to successful candidates, new hire compensation is negotiable.

About Blizzard Entertainment

Develops and publishes interactive video games and services.

Similar jobs

Site Reliability Engineer roles near Irvine, California
4h
Save
Mark Applied
Hide
Site Reliability Engineer Intern (Global SRE) - 2027 Summer
San Jose or Los Angeles or New York City or London or Dublin or Paris or Berlin or Dubai or Jakarta or Seoul or Tokyo
OnsiteInternship, Full Time
TikTok
TikTok: Global short-form video hosting and social media platform.
Currently pursuing a bachelor's degree in computer science or related field; Unix/Linux, IP networking, and Python, Go, C, C++, or Java programming experience required.
Unix/Linux, IP networking, Python, Go, C, C++, Java
4d
Save
Mark Applied
Hide
Senior Site Reliability Engineer - Undersea Dominance
Costa Mesa, California, United States
$166k-$250k/yr OnsiteFull Time
Anduril Industries
Anduril Industries: Defense technology building autonomous military hardware and software.
Advanced Python proficiency; CI/CD, infrastructure-as-code, cloud, containers, monitoring, model registry, parallel computing, and collaboration tool experience; eligible for a U.S. Secret clearance.
Python, GitHub Actions, JFrog Artifactory, Git, CircleCI, Terraform, Ansible, Azure, AWS, Google Cloud Platform (GCP), Docker, Kubernetes, MLflow, Kubeflow, Prometheus, Grafana, CUDA, OpenCL, JIRA, Confluence, C++, Rust, Go
1w
Save
Mark Applied
Hide
Sr. Site Reliability Engineer - Top Secret Clearance (Starlink)
Redmond or Hawthorne
$165k-$230k/yr OnsiteFull Time
SpaceX
SpaceX: Designs and launches advanced rockets and satellite internet constellations.
5+ YOEBachelor's degree in a relevant discipline and 5 years of software development experience, or 7+ years' professional software experience; Linux and active Top Secret or Top Secret/SCI clearance required.
Linux, Kubernetes, Istio, Apache Kafka, Spark, HBase, HDFS, Flink, Python, C#, Java, Scala, Go
1w
Save
Mark Applied
Hide
SRE
Hyderabad or Pasay City or Los Angeles or World Technology Center
OnsiteFull Time
VXI Global Solutions
VXI Global Solutions: Global provider of customer care and business process outsourcing.
Requires observability experience with Prometheus, Grafana, OpenTelemetry, Google Cloud tools, and SolarWinds; telemetry analysis, incident troubleshooting, Python or Bash scripting, and SRE or platform engineering experience preferred.
Prometheus, Grafana, OpenTelemetry, SolarWinds, Google Cloud Platform, Cloud Monitoring, Logging, Trace, Python, Bash, Infrastructure-as-Code, Datadog, New Relic
2w
Save
Mark Applied
Hide
Site Reliability Engineer III
Irvine, California, United States
$133k-$185k/yr OnsiteFull Time
JPMorgan Chase
JPMorgan ChaseNYSE: JPM: Global financial services firm providing banking and investment solutions.
3+ YOE3+ years applied SRE experience, proficiency with SRE principles, one programming language (Python, Java/Spring Boot, .Net), observability, CI/CD, container orchestration, and cloud infrastructure.
Python, Java, Spring Boot, .Net, EKS, EC2, ALB, NLB, Route 53, Terraform, Kubernetes, TLS, mTLS, CloudWatch
3w
Save
Mark Applied
Hide
Senior Site Reliability Engineer
Los Angeles, California, United States
$140k-$180k/yr OnsiteFull Time
K2 Space
K2 Space: Develops high-power satellite platforms for heavy-lift launch vehicles.
5+ YOE5+ years SRE/DevOps experience or BS in CS/IT/STEM, deep cloud (AWS/GCP/Azure), IaC, Kubernetes, Linux, programming (Go/Python), strong security and reliability experience.
Infrastructure-as-Code (IaC), AWS, GCP, Azure, Kubernetes, Terraform, Ansible, Go, Python, Linux
4w
Save
Mark Applied
Hide
Site Reliability Engineer
El Segundo, California, United States
$170k-$195k/yr OnsiteFull Time
Picogrid
Picogrid: Develops hardware and software for autonomous defense systems.
3+ YOE3+ years SRE experience, deep Kubernetes and Terraform/OpenTofu skills, AWS proficiency, observability (Grafana, Prometheus, Loki, OpenTelemetry), incident response, HA databases, and IoT/edge fleet experience.
Grafana, Prometheus, Loki, OpenTelemetry, Terraform, OpenTofu, Kubernetes, AWS, Nebula, WireGuard, Tailscale, Sloth, Pyrra, NVIDIA Jetson
1mo
Save
Mark Applied
Hide
Lead Site Reliability Engineer
Los Angeles, California, United States
$140k-$199k/yr HybridFull Time
Green Dot
Green DotNYSE: GDOT: Provides mobile banking and payment solutions to consumers and businesses.
7+ YOE7+ years in release/reliability engineering, cloud platform experience (AWS/Azure/GCP), automated deployment and observability proficiency, scripting with PowerShell/Bash/Python, excellent troubleshooting and communication skills.
AWS, Azure, GCP, PowerShell, Bash, Python