Apple

Storage Platform Engineer (SRE)

Apple London, England, United Kingdom

Computers and Electronics Manufacturing · 10,001+ employees

5 h ago
sre Senior (5-10 yrs) Full-time United Kingdom
Log in to apply, save this posting, or score it against your profile with AI.

About the role

You will design, build, and maintain large-scale distributed storage systems to support critical Apple Cloud services. The role involves managing on-call rotations, driving technical standards, and collaborating with a globally distributed team to ensure high availability.

What they look for

Distributed storage systems Linux system internals Go Python Java Rust S3 Ceph Gluster NFS Kubernetes Configuration management Capacity planning Disaster recovery Networking protocols Data analysis

Requirements

Candidates must have extensive experience in operating distributed storage systems and proficiency in programming languages such as Go, Java, Python, or Rust. A deep understanding of Linux internals, networking protocols, and storage solutions like S3, Ceph, or NFS is required.

Full description

Apple Cloud infrastructure is BIG. The storage SRE teams of Apple Cloud are building and runningthe next generation distributed storage systems to support Apple’s most critical services. Operating atour scale, across multiple geographically dispersed data centres, and servicing users with vast dataneed presents unique challenges. As a member of Storage SRE at Apple, you'll need to solve theseproblems using your deep understanding of infrastructure, storage, data analysis, programming,teamwork, and expertise in Linux system internals.

Description

We are looking for seasoned software and systems engineers to join the Object Storage SRE team at Apple. You are solution-oriented and have a passion for software delivered as a service to improve reuse, efficiency, and simplicity. Your work will affect hundreds of millions of users and be essential to the success of some of the most visible current and future Apple features.

The role involves understanding the team's priorities; taking ownership of projects or deliverables; designing solutions and building buy-in for those designs; and successful and timely delivery of those designs. Members of our team are expected to be comfortable giving technical feedback to colleagues to assist them in the delivery of their designs, features and projects, as well as driving technical standards across the distributed team in collaboration with key stakeholders.

The team has an on-call rota including the week-ends and the successful candidate should expect to handle alerts and other escalations in order to maintain a high level of availability and functionality for our provided services. The team is globally distributed and cross-timezone meetings are a core feature of how our team collaborates, reaches agreements, and executes to deliver projects.

At Apple Cloud, we run a mix of open source, vendor licensed, and internally developed tools to perform functions such as system configuration management, provisioning, software development & deployment, logging, and monitoring. You'll learn these tools and have opportunities to improve them. We think critically and strive to balance the best solution with the need to get things done for each engineering challenge we face. Good ideas are heard and results are rewarded.

Minimum Qualifications

Experience in building, operating, and scaling distributed storage systems in a private, public,or hybrid cloud environment. The ability to design, author, understand, and release code in languages like Go (preferred),Java, Python, or Rust. Good understanding of block, object, and file storage solutions in Linux (such as LVM, XFS,ext4, S3, Ceph, Gluster, NFS). Understanding of Linux internals, standard networking protocols, and distributed systems. Experience with provisioning, data migration, backup & recovery, at-scale testing, disasterrecovery, and capacity planning.

Preferred Qualifications

Awareness of best practices for deployment of storage systems - implication of physical and virtual deployment models to change management, failure domains, hardware lifecycle management, etc. Acute drive to automate manual operations and to improve them through repeated iteration. Hands-on experience managing large numbers of diverse systems with configuration management or software delivery platforms (such as Puppet, Chef, Ansible, and Spinnaker). Experience with deploying, supporting and monitoring new and existing services, platforms, and application stacks. Familiarity with microservices architecture and container orchestration with Kubernetes. Familiarity with relational & non-relational databases (such as Cassandra, Postgres, & RocksDB). Familiarity with microservices architecture and container orchestration with Kubernetes. Familiarity with relational & non-relational databases (such as Cassandra, Postgres, & RocksDB).

Similar roles