Apple

Senior Data Services SRE - ASE Databases

Apple Shanghai, Shanghai, China

Computers and Electronics Manufacturing · 10,001+ employees

8 h ago
sre Senior (5-10 yrs) Full-time China
Create a free account to apply — email only, no card. You can also save this posting or score it against your profile with AI.

About the role

You will own the reliability, performance, and scalability of critical data service platforms at Apple scale. This involves developing maintenance automation, backup services, and monitoring tools while contributing to open-source projects.

What they look for

Distributed systems Database systems SRE Java Go Python Apache Cassandra Apache Kafka Apache Solr Redis Kubernetes Performance engineering Monitoring Incident management Terraform Cloud infrastructure

Requirements

Candidates must have demonstrated expertise in distributed systems or database engineering and proficiency in Java, Go, or Python. A degree in Computer Science or equivalent experience is required, along with a strong understanding of SRE concepts and distributed systems fundamentals.

Full description

Apple's Services Engineering organization (ASE) is seeking experienced distributed systems and database engineers to join our Databases SRE organization in Shanghai.

The Databases SRE organization runs the data service technologies that Apple's internet services are built on, Apache Cassandra, Apache Kafka, Apache Solr, Redis/Valkey, and coordination systems such as etcd and Apache ZooKeeper. Each of these is owned by a dedicated SRE team, and this posting covers engineering roles across all of them. You will join one of these teams, matched to your background and to where the Shanghai site needs depth; you are not expected to arrive with experience in all of these technologies.

Description

ASE Databases SRE teams develop applications and tooling that are safe, reliable, scalable, and fast. This work requires an innovative spirit and an extraordinary degree of care and rigor in engineering. Team members contribute to all major components of our data service deployment infrastructure, including maintenance automation, backup services, monitoring and alerting tooling and dashboards, deployment architecture, and upstream contributions to the open source projects we run, focused on stability, performance, and scaling.

The team's work is deployed at massive scale, serving millions of queries per second over hundreds of petabytes of data across our data centers worldwide. It also has big impact, forming the platform upon which iCloud and many other internet services at Apple are built. In ASE, your work will benefit hundreds of millions of users and is critical to the success of some of the most visible current and future Apple features.

Day to day, you will own the reliability and performance of one of these platforms at Apple scale. The problems are shared across our teams even when the systems differ: replication and consistency, capacity and cost, safe rollouts across thousands of nodes, multi-datacenter failure domains, and the long tail of performance. Engineers regularly work across team boundaries on tooling and infrastructure common to all of our data services, so depth in one system and curiosity about the others matters more than experience with any specific technology.

Minimum Qualifications

Demonstrated expertise developing distributed systems, database systems, storage engines, streaming or messaging systems, search platforms, coordination/consensus services, or performance engineering Depth in at least one data service technology, and the curiosity to learn others (the specific system matters less than the distributed systems fundamentals behind it) Understanding of core SRE concepts: monitoring, alerting, and incident management Understanding of distributed systems concepts (consistency models, replication and consensus protocols, isolation levels, crash and recovery semantics) Proficient in one or more of: modern Java, Go (golang), Python Operating systems concepts (process scheduling, disk and network I/O, performance) Strong written and verbal communication skills, and the ability to work effectively with a geographically distributed team Education: BS or MS in Computer Science / related fields or equivalent work experience

Preferred Qualifications

Experience developing critical internet services and/or platform infrastructure Performance engineering (design concepts, profile-guided optimization) Service management and migrations across bare metal, virtualized (EC2), and containerized (Kubernetes) platforms Experience managing services run on Kubernetes Experience with EC2, EBS, and Terraform Fundamentals of system-level hardware and networking components (storage devices and controllers, network interfaces, CPU and memory layout in server-class systems) Datacenter architecture (networking topologies, host placement strategies, and failure modes); design of multi-datacenter systems; failure domains; and wide-area networking Experience contributing to open source distributed systems projects Prior experience with development or maintenance of distributed databases, streaming systems, or storage systems

Similar roles