Senior Data Services SRE - ASE Databases
Apple Shanghai, Shanghai, China
Computers and Electronics Manufacturing · 10,001+ employees
About the role
You will own the reliability, performance, and scalability of critical data service platforms at Apple scale. This involves developing maintenance automation, backup services, and monitoring tools while contributing to open-source projects.
What they look for
Requirements
Candidates must have demonstrated expertise in distributed systems or database engineering and proficiency in Java, Go, or Python. A degree in Computer Science or equivalent experience is required, along with a strong understanding of SRE concepts and distributed systems fundamentals.
Full description
Apple's Services Engineering organization (ASE) is seeking experienced distributed systems and database engineers to join our Databases SRE organization in Shanghai.
The Databases SRE organization runs the data service technologies that Apple's internet services are built on, Apache Cassandra, Apache Kafka, Apache Solr, Redis/Valkey, and coordination systems such as etcd and Apache ZooKeeper. Each of these is owned by a dedicated SRE team, and this posting covers engineering roles across all of them. You will join one of these teams, matched to your background and to where the Shanghai site needs depth; you are not expected to arrive with experience in all of these technologies.
Description
ASE Databases SRE teams develop applications and tooling that are safe, reliable, scalable, and fast. This work requires an innovative spirit and an extraordinary degree of care and rigor in engineering. Team members contribute to all major components of our data service deployment infrastructure, including maintenance automation, backup services, monitoring and alerting tooling and dashboards, deployment architecture, and upstream contributions to the open source projects we run, focused on stability, performance, and scaling.
The team's work is deployed at massive scale, serving millions of queries per second over hundreds of petabytes of data across our data centers worldwide. It also has big impact, forming the platform upon which iCloud and many other internet services at Apple are built. In ASE, your work will benefit hundreds of millions of users and is critical to the success of some of the most visible current and future Apple features.
Day to day, you will own the reliability and performance of one of these platforms at Apple scale. The problems are shared across our teams even when the systems differ: replication and consistency, capacity and cost, safe rollouts across thousands of nodes, multi-datacenter failure domains, and the long tail of performance. Engineers regularly work across team boundaries on tooling and infrastructure common to all of our data services, so depth in one system and curiosity about the others matters more than experience with any specific technology.
Minimum Qualifications
Demonstrated expertise developing distributed systems, database systems, storage engines, streaming or messaging systems, search platforms, coordination/consensus services, or performance engineering Depth in at least one data service technology, and the curiosity to learn others (the specific system matters less than the distributed systems fundamentals behind it) Understanding of core SRE concepts: monitoring, alerting, and incident management Understanding of distributed systems concepts (consistency models, replication and consensus protocols, isolation levels, crash and recovery semantics) Proficient in one or more of: modern Java, Go (golang), Python Operating systems concepts (process scheduling, disk and network I/O, performance) Strong written and verbal communication skills, and the ability to work effectively with a geographically distributed team Education: BS or MS in Computer Science / related fields or equivalent work experience
Preferred Qualifications
Experience developing critical internet services and/or platform infrastructure Performance engineering (design concepts, profile-guided optimization) Service management and migrations across bare metal, virtualized (EC2), and containerized (Kubernetes) platforms Experience managing services run on Kubernetes Experience with EC2, EBS, and Terraform Fundamentals of system-level hardware and networking components (storage devices and controllers, network interfaces, CPU and memory layout in server-class systems) Datacenter architecture (networking topologies, host placement strategies, and failure modes); design of multi-datacenter systems; failure domains; and wide-area networking Experience contributing to open source distributed systems projects Prior experience with development or maintenance of distributed databases, streaming systems, or storage systems
Similar roles
-
Senior Software Development Engineer (SRE)
Renesas Electronics Belgrade, Central Serbia, Serbia
-
Sr. Staff Site Reliability Engineer
Zscaler Bengaluru, Karnataka, India
-
Site Reliability Engineer
trivago Düsseldorf, North Rhine-Westphalia, Germany
-
Site Reliability Engineer-1
Digantara Bengaluru, Karnataka, India
-
Senior Staff Site Reliability Engineer – Compute Platform
NVIDIA Bengaluru, Karnataka, India
-
Site Reliability Engineer – Quantum Computing (f/m/d)
eleQtron GmbH Hamburg, Germany