Apple
Posted 4mo ago

Site Reliability Engineer, Enterprise Technology Services

Apple
Sunnyvale or United States
OnsiteFull Time
Responsibilities
  • designing tools
  • building platforms
  • monitoring systems
Requirements
  • 5+ years SRE experience with large-scale distributed systems
  • Java
  • CS degree or equivalent
  • Strong Open Source data processing
  • Observability stack expertise
  • Troubleshooting
Technical tools mentioned
PrometheusGrafanaDatadogOpenTelemetryELKKubernetesHelmCI/CDGitAnsible

Job description

At Apple, groundbreaking ideas quickly transform into extraordinary products and services that delight millions worldwide. If you’re passionate about engineering and operating robust, large-scale systems, imagine the impact you could make.

The Identity Management Services (IdMS) SRE team is seeking a Service Reliability Engineer (SRE) to design, build tools for, and support our critical platform services. We’re looking for someone with strong software development skills, deep systems expertise, and a solid understanding of SRE principles, ready to ensure operational precision at Apple’s immense scale. Your work will be pivotal in powering services across Apple, partnering with engineering teams to deliver seamless experiences.

Description

This role involves managing one of the largest Identity Management Platform services for a vast customer base across various devices and services. Key responsibilities include overseeing critical services such as device provisioning, authentication, token management, and security. A primary objective is ensuring the high availability and reliability of the system to facilitate critical authentication and authorization transactions, user provisioning, purchases, subscriptions, and account lifecycle management (creation, management, and recovery). This also entails maintaining platform security by blocking and rate-limiting fraud traffic at the perimeter, and ensuring high data consistency and replication across multiple data centers through custom mechanisms. The role covers managing infrastructure, capacity planning, disaster recovery, and auto-failover mechanisms. It also involves monitoring infrastructure and application services, driving incident management for internal and external stakeholders, and defining system and functional observability. Furthermore, this position helps teams overcome system bottlenecks and architectural challenges for efficiency improvements, ensures systems are compliant with industry standards and pass critical audits, and drives automation solutions for large-scale platform service needs. Advanced responsibilities include alert engineering, anomaly detection with Machine Learning tools, and adapting to Generative AI enhancements. Investigating device-related issues by debugging relevant logs is also part of the role, alongside managing the full system lifecycle, including configuration and code deployment in user acceptance test and production environments.

Minimum Qualifications

  • 5+ years of experience in Site Reliability Engineering with a strong focus on building, scaling, and operating large-scale distributed platform services, and Java.
  • BS degree in computer science or equivalent field with 7+ years of experience or MS degree in computer science or equivalent field with 5+ years of experience.
  • Strong technical grasp and experience working on Open Source technologies designed for large-scale data processing.
  • Experience designing, analyzing, and troubleshooting distributed systems.
  • Proficiency in at least one programming or scripting language (Python, Java, Go, Bash, Ansible, or similar).
  • Experience designing observability stacks (Prometheus, Grafana, Datadog, OpenTelemetry, ELK, etc.).
  • Excellent troubleshooting and problem-solving skills.

Preferred Qualifications

  • Observability & SRE Principles: Experience with monitoring and logging tools (e.g., Prometheus, Splunk, Grafana, OpenTelemetry) and a strong understanding of SRE principles, including observability, error budgeting, and service reliability metrics (SLA, SLO, SLI).
  • CI/CD & Automation: Proficiency with CI/CD, Release Engineering, DevOps practices, and source control (Git). Experience designing and implementing CI/CD pipelines and Infrastructure as Code (Helm, CRD).
  • Programming & Data Systems: Strong programming skills in languages like Java, Python, Go, etc. Experience with various databases (Relational, NoSQL, OLAP) and event-driven architectures (Kafka, RabbitMQ).
  • Reliability & Operations: Experience with on-call, including incident/problem management (PIR, RCA) and a strong sense of ownership for system reliability.
  • Security & Compliance: Understanding of security standards, policies, cryptography, and authentication (OAuth, SAML, SSO). Knowledge of Governance and Compliance.
  • Innovation & Collaboration: Passion for designing reliable systems, advocating for automation, and a desire to collaborate effectively. Experience leveraging ML/GenAI for operational efficiency is a plus.
  • Certification: Cybersecurity certification will be an added advantage.
  • Education: Bachelor’s or Master’s degree in Computer Science or equivalent practical experience.

About Apple

Designs and sells consumer electronics, software, and online services.

Similar jobs

Site Reliability Engineer roles near Sunnyvale, California
2h
Save
Mark Applied
Hide
Site Reliability Graduate (Data Infrastructure) - 2027 Start
San Jose, California, United States
OnsiteFull Time
ByteDance
ByteDance: Developing AI-driven content platforms and mobile applications.
Bachelor's degree in computer science, computer engineering, or related field required; master's preferred. Requires scripting, Linux, and networking knowledge. Docker, Kubernetes, data stores, observability, and data center experience preferred.
Kubernetes, Redis, MySQL, Message Queue, Python, Go, Bash, Linux, Docker, PostgreSQL, Prometheus, Grafana, ELK Stack
12h
Save
Mark Applied
Hide
Site Reliability Engineer (US - Pacific time)
San Francisco or United States
$271k-$296k/yr RemoteFull Time
PostHog
PostHog: All-in-one product analytics and developer tools platform
Requires hands-on production Kubernetes and AWS experience, Terraform or Terragrunt automation, Linux systems knowledge, stateful infrastructure experience, production debugging, and end-to-end on-call ownership.
Amazon Web Services (AWS), Kubernetes, Amazon Elastic Kubernetes Service (EKS), Karpenter, Cilium, ArgoCD, Terraform, Terragrunt, GitHub Actions, Linux, Cloudflare, Microsoft SharePoint
21h
Save
Mark Applied
Hide
Senior Site reliability Engineer (Linux) IRC302310
San Jose, California, United States
$130k-$140k/yr RemoteFull Time
GlobalLogic
GlobalLogic: Digital product engineering and software development services provider.
7+ YOERequires 7+ years of Linux production experience, automation skills, debugging expertise, cloud and on-premises experience, and a bachelor's or master's degree in a related field.
AWS, Linux, Python, Ruby, Ansible, Docker, Kubernetes, AMI, C, Go
1d
Save
Mark Applied
Hide
Senior Site Reliability Engineer, BCM - DGX Cloud
Santa Clara or United States
$168k-$334k/yr HybridFull Time
NVIDIA
NVIDIANASDAQ: NVDA: Designs GPU-accelerated computing and artificial intelligence hardware.
8+ YOEBachelor's degree or equivalent in computer science or related field, 8+ years in site reliability engineering or software development, Python, Linux, networking, and cluster operations experience.
Python, Linux, C++, Kubernetes, Slurm, InfiniBand, Spectrum-X, BCM, Bright Cluster Manager, Base Command Manager
3d
Save
Mark Applied
Hide
Site Reliability Engineer, US Gov
Denver or Arvada or San Francisco or Nashville or Santa Fe or New Orleans or San Diego or Bozeman or United States
$160k-$200k/yr HybridFull Time
Quindar
Quindar: Cloud-native software for automated satellite mission operations.
3+ YOERequires a bachelor's degree, 3+ years of SRE or infrastructure experience, U.S. citizenship, Secret clearance or higher, and expertise in Kubernetes, AWS, Python, Terraform, networking, and CI/CD.
AWS GovCloud, AWS C2E, Kubernetes, AWS EKS, Rancher, Grafana LGTM, Datadog, Python, Terraform, VPN, NLB, ALB, HTTPS, TLS, VPC peering, CDN, GitLab Workflows, Unix, Linux, Auth0, Keycloak, AWS IAM, Git
3d
Save
Mark Applied
Hide
Senior Site Reliability Engineer
San Francisco or Bellevue
$240k-$356k/yr HybridFull Time
Lambda
Lambda: Provides high-performance GPU cloud infrastructure for AI development.
7+ YOERequires 7+ years in SRE, HPC engineering, DevOps, or similar; expertise in AI infrastructure, Linux, distributed systems, networking, Python, Go, monitoring, and automation tools.
Linux, Ansible, Terraform, Python, Go, Prometheus, Grafana, ClickHouse, PyTorch, TensorFlow, DeepSpeed, MLPerf, Docker, Kubernetes, InfiniBand, RoCE, NCCL, GPU-direct, CLOS, 100GbE, Ethernet, SOC 2, ISO 27001
4d
Save
Mark Applied
Hide
Site Reliability Engineer - rednote
Palo Alto, California, United States
OnsiteFull Time
Rednote
Rednote: A lifestyle-focused social media and e-commerce discovery platform.
Experience with large-scale reliability, high-availability architecture, incident response, cross-region disaster recovery, Linux, networking, middleware, cloud-native infrastructure, automation, and Python, Go, or Java.
Linux, MySQL, Redis, Kafka, Kubernetes, Service Mesh, Python, Go, Java
4d
Save
Mark Applied
Hide
K8 Site Reliability SME
San Jose or Austin
RemoteFull Time
Bitdeer
BitdeerNASDAQ: BTDR: Operates cryptocurrency mining and high-performance computing data centers.
5+ YOERequires 5+ years of Kubernetes operations, 2+ years managing GPU workloads, Terraform, Helm, GitOps, SRE practices, monitoring, Go or Python, and multi-tenant platform experience.
Kubernetes, Nvidia GPU operator, Terraform, Helm, ArgoCD, Flux, Prometheus, Grafana, Alertmanager, PagerDuty, Go, Python, Slurm, Ray, Kubeflow, Ironic, MAAS, GitOps