Apple

Site Reliability Engineer (SRE), London

Apple London, England, United Kingdom

Computers and Electronics Manufacturing · 10,001+ employees

5 h ago
sre Senior (5-10 yrs) Full-time United Kingdom
Create a free account to apply — email only, no card. You can also save this posting or score it against your profile with AI.

About the role

The Site Reliability Engineer will be responsible for the availability, automation, and operational excellence of critical systems supporting Apple's Private Cloud Compute. You will contribute to building and scaling secure, privacy-preserving cloud infrastructure while solving complex technical challenges.

What they look for

Kubernetes GKE EKS Distributed Systems GCP AWS Automation Java Swift Python TypeScript Nginx Envoy Prometheus Docker Linux

Requirements

Candidates must have hands-on experience operating managed Kubernetes in public clouds and scaling distributed systems. Proficiency in a high-level programming language such as Java, Swift, Python, or TypeScript is required, along with experience in infrastructure automation and disaster recovery.

Full description

People at Apple don’t just build products — they craft experiences our customers love and depend on. Apple Services Engineering (ASE) builds and supports the systems that make many of these daily experiences possible. If you’ve used Apple products, you’ve likely interacted with us. Private Cloud Compute (PCC) represents a groundbreaking approach to cloud intelligence, extending the security and privacy of Apple devices into the cloud to unlock even more intelligence for our users. This SRE team is responsible for the availability and automation of the critical systems and services that enable PCC to deliver cloud intelligence without compromising user privacy. If you're passionate about building the future of privacy-preserving cloud infrastructure at scale, this is the opportunity for you!

Description

We're looking for a hardworking and passionate SRE Engineer to join this amazing team. You will be an accomplished builder and problem-solver, eager to tackle challenging technical problems. You have a deep understanding of SRE principles and the expertise required to operate services at Apple scale with a high degree of operational excellence. This role will allow you to directly contribute to shaping the future of how we build and run our services on a global scale. You will possess strong technical skills to dive deep into complex systems while also understanding and contributing to higher-level business and product goals. We seek high-quality engineers with a diverse set of experiences and skill sets. Our customers count on us to provide extraordinary availability, scalability, and security for services. If you’d like to positively influence millions of customers’ experience of Apple through your technical contributions, this is the job for you.

Minimum Qualifications

In depth hands-on experience operating managed Kubernetes (GKE and/or EKS) in a public cloud, with experience scaling distributed systems. Experience across multiple public clouds (GCP and AWS) strongly preferred. Strong experience with deploying, supporting and supervising new and existing services, platforms and application stacks Experience with scale testing, disaster recovery, and capacity planning Passion for eliminating repetitive manual processes using automation to improve them through repeated iteration Confirmed ability to write programs using a high-level programming language like: Java, Swift, Python, or TypeScript Proclivity towards efficient programming emphasizing improvement via complexity analysis. Experience with Nginx, Envoy, Prometheus, and/or Docker.

Preferred Qualifications

Understanding of standard networking protocols and components such as: HTTP, DNS, ECMP, TCP/IP, ICMP, the OSI Model, Subnetting and Load Balancing strategies. Understanding of the Linux Operating System, including Kernel, Memory, Process, Threads, Static / Shared Libraries, IPC, Signals. Experience with Infrastructure-as-Code and config-as-code tooling such as Pulumi, Terraform, or Pkl. Experience with fleet/cluster lifecycle management, node provisioning, and hardware-adjacent reliability (e.g., GPU health, capacity management) at scale. Experience building and operating CI/CD pipelines for cloud infrastructure.

Similar roles