Apple

Lead SRE

Apple Shanghai, Shanghai, China

Computers and Electronics Manufacturing · 10,001+ employees

10 h ago
sre Principal (10+ yrs) Full-time China
Create a free account to apply — email only, no card. You can also save this posting or score it against your profile with AI.

About the role

The Lead SRE will provide technical leadership for the Apple Data Platform Compute team to ensure infrastructure reliability and smooth 24x7 operations. Responsibilities include managing core infrastructure, supporting parallel migrations, and driving automation and process management.

What they look for

Site Reliability Engineering Technical Leadership Kubernetes AWS GCP Alibaba Cloud Linux Networking Containers Infrastructure Management Automation Incident Management Capacity Planning Disaster Recovery Data Analytics Performance Tuning

Requirements

Candidates must have 12+ years of SRE experience with at least 5 years in leadership roles and proficiency in managing Kubernetes clusters on major cloud platforms. Advanced knowledge of Linux, networking, and container technologies is required, along with a proven history of project delivery.

Full description

Apple Service Engineering (ASE) teams build and scale the platforms and infrastructure behind many of Apple's services (such as iCloud, iTunes, Siri, and Maps). We are the foundation on which Apple's software developers build the products that our customers love.

Description

We are looking for a passionate and dedicated Senior Site Reliability Engineer to provide technical leadership on our team to help ensure our customers have the highest quality Apple Services experience. The Apple Data Platform (ADP) Compute SRE team is responsible for the core infrastructure, including our legacy bare-metal platforms and modern cloud based infrastructure stack. We partner with both peer SRE teams and several of our world-class software and product engineering teams to support infrastructure reliability, multi-year parallel migrations for Apple properties, as well as the automation, tooling, incident, and process management necessary to ensure smooth 24x7 operations for ADP customers.

Minimum Qualifications

12+ years of experience in Site Reliability Engineering, specifically managing infrastructure and services at scale 5+ years of experience in management or technical leadership roles 5+ years of proficiency in running applications and managing Kubernetes clusters on Alibaba Cloud, AWS, or GCP Proven history of end-to-end project management and delivery Demonstrable programming skills for developing software/tools and leading code reviews Advanced knowledge of Linux, Networking, and Containers

Preferred Qualifications

15+ years of experience in SRE or related work managing infrastructure at scale Proficiency with the architecture, deployment, performance tuning, and troubleshooting of open source data analytics or governance technologies such as Spark, Flink, Iceberg, Trino, and/or Druid. Experience with scale testing, disaster recovery, and capacity planning Ability to define the technical roadmap for infrastructure and drive cross-functional alignment on architectural standards and best practices

Similar roles