Apple
Posted 2w ago

Site Reliability Engineer, Apple Data Platform / Multi-Cloud Infrastructure

Apple
Austin, Texas, United States
OnsiteFull Time
Responsibilities
  • managing incidents
  • supporting teams
  • operating platform
Requirements
  • Manage and operate a massive multi-cloud data platform
  • Run incident response
  • Provide hands-on support to internal teams, and partner with developers to keep services reliable across AWS, GCP, and on-prem Kubernetes
Technical tools mentioned
SparkFlinkAirflowRayNotebooksKubernetesAWSGCP

Job description

The Apple Services Engineering team (ASE) is one of the most exciting examples of Apple's long-held passion for combining art and technology. These are the people who power the App Store, Apple TV, Apple Music, Apple Podcasts, and Apple Books — at extensive scale, meeting high expectations to deliver a huge variety of entertainment in over 35 languages to more than 150 countries.

Within ASE, the Apple Data Platform SRE team keeps a massive, multi-cloud platform running for thousands of internal engineers building the next generation of data and AI products at Apple. We sit at the intersection of infrastructure, automation, and customer success — running incident response, providing hands-on support to internal teams, and partnering with developers to make cutting-edge services like Spark, Flink, Airflow, Ray, Notebooks, and LLM-based agent platforms reliable at scale across AWS, GCP, and on-premise Kubernetes.

About Apple

Designs and sells consumer electronics, software, and online services.

Similar jobs

Site Reliability Engineer roles near Austin, Texas
1d
Save
Mark Applied
Hide
K8 Site Reliability SME
San Jose or Austin
RemoteFull Time
Bitdeer
BitdeerNASDAQ: BTDR: Operates cryptocurrency mining and high-performance computing data centers.
5+ YOERequires 5+ years of Kubernetes operations, 2+ years managing GPU workloads, Terraform, Helm, GitOps, SRE practices, monitoring, Go or Python, and multi-tenant platform experience.
Kubernetes, Nvidia GPU operator, Terraform, Helm, ArgoCD, Flux, Prometheus, Grafana, Alertmanager, PagerDuty, Go, Python, Slurm, Ray, Kubeflow, Ironic, MAAS, GitOps
2d
Save
Mark Applied
Hide
Cleared Senior Site Reliability Engineer
Austin, Texas, United States
$80k-$210k/yr OnsiteFull Time
Gallatin
Gallatin: Develops AI software for military logistics and supply chains.
3+ YOERequires active Secret clearance, 3–5 years in SRE, DevOps, or production infrastructure, Linux, networking, cloud infrastructure, infrastructure-as-code, containers, CI/CD, monitoring, and incident response experience.
Linux, AWS, Azure, Terraform, Ansible, Kubernetes, Docker, Prometheus, Grafana, Datadog, ELK, Microsoft Azure Government, Microsoft Azure Government Secret, Microsoft Azure Government IL5, Microsoft Azure Government IL6
2d
Save
Mark Applied
Hide
Site Reliability Engineer Spring Co-op 2027
Lowell or Durham or San Jose or Austin
$76k-$166k/yr HybridMultiple Commitments Available
IBM
IBMNew York Stock Exchange: IBM: Global technology providing enterprise software, cloud, and consulting.
Actively enrolled in a bachelor's program, available for a 16-week full-time co-op, and knowledgeable in Linux, monitoring, troubleshooting, automation, scripting, cloud platforms, and production support.
Linux, Python, Go, Bash, IBM Cloud, AWS, Microsoft Azure, Google Cloud Platform, Kubernetes, OpenShift, Ansible, Terraform, Jenkins, IBM Continuous Delivery, ArgoCD, Instana, New Relic, Grafana, Prometheus, PostgreSQL, CouchDB, Redis, Kafka, Spark, SQL, NoSQL, CI/CD
3d
Save
Mark Applied
Hide
Site Reliability Engineer
Austin, Texas, United States
$112k-$191k/yr HybridFull Time
Thales
ThalesEuronext Paris: HO: Develops electronics and digital systems for aerospace and defense.
5+ YOERequires 5+ years in cloud, web, or CDN infrastructure; Python and Go; Linux, networking, DevOps, distributed systems, and on-call experience; BS/MS in computer science, engineering, or equivalent experience.
Python, Go, C, C++, Linux, TCP, UDP, DNS, TLS/SSL, HTTP, BGP, Ansible, Saltstack, GitLab, Jenkins, Git, Prometheus, Grafana, NoSQL, RDBMS, Redis, Elasticsearch, Kafka, Docker, Kubernetes
5d
Save
Mark Applied
Hide
Site Reliability Engineer
Austin, Texas, United States
HybridFull Time
Thales
ThalesEuronext Paris: HO: Designs and manufactures electronic systems for aerospace and defense.
5+ YOEEngineer or equivalent with at least 5 years of experience, Java development, public cloud, containers, microservices, CI/CD, automation, monitoring, and observability. U.S. or dual citizenship required.
Terraform, Ansible, Kubernetes, GitLab, Datadog, Java, GCP, AWS, Docker, Jenkins, Helm, NoSQL
5d
Save
Mark Applied
Hide
AMHS Site Reliability Engineer
Taylor, Texas, United States
$90k-$115k/yr OnsiteFull Time
Samsung Electronics
Samsung ElectronicsKorea Exchange: 005930: Develops and manufactures consumer electronics, semiconductors, and mobile devices.
3+ YOEBachelor’s degree in a technical field, 3+ years in semiconductor-related industries, SRE/software/DevOps experience, backend data engineering, Linux/Unix administration, observability, and strong English skills.
Python, Java, Prometheus, Grafana, Docker, Kubernetes, RabbitMQ, Redis, Linux, Unix, ETL, NoSQL, IPC, RPC, SECS/GEM, MES
1w
Save
Mark Applied
Hide
Sr. Site Reliability Engineer - Core Platform & Embedded Reliability (Hybrid)
New York City or Austin or Sunnyvale or Redmond
$140k-$215k/yr HybridFull Time
CrowdStrike
CrowdStrikeNASDAQ: CRWD: Provides cloud-native endpoint protection and cybersecurity services.
10+ YOE10+ years building distributed systems, 5+ years developing SaaS microservices, expert programming skills, distributed-systems expertise, architectural leadership, and a Computer Science degree or equivalent experience.
Go, Java, Scala, Kotlin, Python, Node.js, Kubernetes, AWS, Cassandra, Kafka, Elasticsearch, OpenSearch, Google Cloud Platform (GCP), Oracle Cloud Infrastructure (OCI), GitHub, Stack Overflow
1w
Save
Mark Applied
Hide
Senior Site Reliability Engineer
Austin, Texas, United States
HybridFull Time
TeamViewer
TeamViewerFrankfurt Stock Exchange: TMV: Provides remote connectivity and digital workplace software solutions.
5+ YOEDegree in computer science, software engineering, IT, or equivalent experience; 5+ years in SRE, DevOps, or software development; Azure, IaC, automation, containers, monitoring, databases, security, and scripting expertise.
Microsoft Azure, Kubernetes, GitOps, Terraform, Argo CD, PowerShell, Docker, MS SQL, Postgres, Datadog, Grafana, Prometheus