Site Reliability Engineering (SRE) Manager, Apple Maps
Cupertino, California, United States
OnsiteFull Time
AppleNASDAQ: AAPL: Designs and sells consumer electronics, software, and online services.
Build, manage, and deliver highly available, automated infrastructure for Apple Maps at global scale; focus on reliability, scalability, and operational excellence.
Lambda: Provides high-performance GPU cloud infrastructure for AI development.
8+ YOE8+ years in incident management/SRE/infrastructure operations; experience with large-scale distributed infrastructure, data center operations, GPU clusters, networking, cloud platforms; incident frameworks (ITIL/SRE); strong leadership, communication, and stakeholder management.
PagerDuty, ServiceNow, Jira, Datadog, Prometheus, Grafana, Incident command system (ICS)
Lambda: Provides high-performance GPU cloud infrastructure for AI development.
Experience with system design, scalable cloud infrastructure, configuration management, programming in Python or Go, distributed systems, automation, documentation, and cross-functional collaboration.
AWS, GCP, Azure, Chef, Ansible, Terraform, GitHub Actions, Python, Go
Kody: An agentic commerce platform providing integrated in-person payment solutions.
Deep AWS and GitHub experience, strong monitoring/logging and scripting skills, incident management ownership, and absolute fluency in Mandarin and English.
ZoomNasdaq: ZM: Provides a cloud-based platform for video, voice, and collaboration.
5+ YOE3+ Mgmt5+ years SRE/DevOps experience,3+ years managing technical teams,BS/MS or equivalent,expertise with Kubernetes,Terraform,cloud providers,CI/CD and incident response,programming and scripting skills.
Bloom EnergyNYSE: BE: Manufactures solid oxide fuel cell systems for onsite power.
10+ YOE3+ MgmtBachelor’s degree and 10+ years in cloud, infrastructure, platform engineering, or DevOps, including 3+ years in senior technical leadership. Requires AWS, networking, IaC, CI/CD, Kubernetes, observability, security, and SRE expertise.
BitdeerNASDAQ: BTDR: Operates cryptocurrency mining and high-performance computing data centers.
5+ YOERequires 5+ years of Kubernetes operations, 2+ years managing GPU workloads, Terraform, Helm, GitOps, SRE practices, monitoring, Go or Python, and multi-tenant platform experience.
ZoomNasdaq: ZM: Provides a cloud-based platform for video, voice, and collaboration.
15+ YOE5+ Mgmt15+ years leading IT infrastructure and operations with 5+ years in senior management; experience managing large global teams, IAM, SRE/automation, AI/ML in support environments, XaaS and on‑prem architectures, and compliance frameworks (NIST, SOX, SOC).
ByteDance: Developing AI-driven content platforms and mobile applications.
Bachelor's in related field and strong experience with large-scale Linux host management, core data-center services (DNS, NTP, DHCP, NAT, APT, Kerberos), DevOps tooling, SRE practices, and troubleshooting.
Senior Site Reliability Engineer - Managed Kubernetes
San Francisco or San Jose or Bellevue
$240k-$356k/yrHybridFull Time
Lambda: Provides high-performance GPU cloud infrastructure for AI development.
6+ YOERequires 6+ years in SRE or operations, deep Linux and production Kubernetes expertise, strong Go and Python skills, GitOps, Helm, observability, CI/CD, and Kubernetes provisioning experience.
Tech Lead Cloud Site Reliability Engineer - DCS Cloud
San Jose, California, United States
OnsiteFull Time
ByteDance: Developing AI-driven content platforms and mobile applications.
5+ YOE5+ years SRE/DevOps/Linux operations experience, bachelor’s in CS or related field, proficiency in Go/Python/C++, strong troubleshooting, monitoring, and reliability practices.
Slalom: Provides business and technology consulting and software engineering services.
Senior leader with deep technology delivery and leadership experience in product engineering, cloud modernization, AI-accelerated engineering, SRE/operations, and go-to-market capability development.
Lambda: Provides high-performance GPU cloud infrastructure for AI development.
10+ YOE10+ years in software/platform engineering or SRE; 5+ years Kubernetes at scale; strong Go and Python; deep Kubernetes internals; GPU orchestration; multi-tenant infrastructure; distributed systems; observability at scale; Linux networking; IaC and GitOps.
Tech Lead Site Reliability Engineer, TikTok Generalized Arch USTO
San Jose, California, United States
$245k-$450k/yrOnsiteFull Time
TikTok: Global short-form video hosting and social media platform.
5+ YOEBachelor's in CS or related, strong CS foundation, Linux and storage/network knowledge, proficiency in Python/Go/Java/PHP/C/C++, strong problem solving and communication; 5+ years SRE/cloud experience preferred.
Director, Enterprise IT Infrastructure (CA, US, 95110)
San Jose, California, United States
$191k-$280k/yrOnsiteFull Time
QuantumScapeNYSE: QS: Develops next-generation solid-state batteries for electric vehicles.
15+ YOE5+ Mgmt15+ years IT experience with 5+ years leading enterprise infrastructure; Bachelor's in CS/IT/Engineering; hands-on expertise with GCP/Azure, Kubernetes (GKE), networking, identity, endpoints, SRE, automation, and OT/IT integration; strong leadership and cross-functional communication.