Apple

ASE Compute - Senior SRE Software Engineer

Apple · San Francisco, California, United States

Computers and Electronics Manufacturing · 10,001+ employees

6 h ago
Principal (10+ yrs) Full-time United States
Log in to apply, save this posting, or score it against your profile with AI.

About the role

You will own the technical architecture and reliability of Kubernetes internals at global scale while mentoring engineers and driving upstream community contributions. The role involves leading incident response, defining SLOs, and building AI-assisted tooling to automate operational workflows.

What they look for

Kubernetes Golang SRE System Architecture Infrastructure Automation Distributed Systems Cloud-native Prometheus Thanos Grafana Controller-runtime Etcd Apiserver Kubelet Incident Response SLO Management

Requirements

Candidates must have a Bachelor's degree in Computer Science or related field and deep experience operating core Kubernetes components in production. Expert-level proficiency in Golang and a proven track record of technical leadership and end-to-end system ownership are required.

Full description

People at Apple don't just build products, they craft the kind of experience that has revolutionized entire industries. The diverse collection of our people and their ideas inspire innovation in everything we do. Imagine what you could do here! Join Apple, and help us leave the world better than we found it.

The Apple Service Engineering (ASE) team builds and provides systems and infrastructure that power Apple's services (such as iCloud, Apple Music, Apple Intelligence, and Maps). We are the foundation on which Apple's software developers build the products that our customers love. Our services have to scale globally, stay highly available, and "just work." If you love designing, engineering, and running systems and infrastructure that will help millions of customers, then this is the place for you!

Description

The ASE Compute team is looking for a senior SRE software engineer to own the technical direction of the Kubernetes internals that Apple's services run on. You will set the architecture for our controllers and namespace management infrastructure, strengthen the reliability of our Kubernetes services, and engage with the upstream community to drive Apple's requirements. You will write the hardest parts yourself, and raise what the rest of the team can build through design review, mentoring, and the tools you build. Service teams across Apple will come to you as the technical authority on what the platform can do and where it is going. The role also offers room to build AI-assisted tooling that accelerates triage, operational workflows, and infrastructure automation for the whole team.

Responsibilities

* Own the architecture and technical direction for our controllers and namespace management infrastructure, from design through production operation at global scale. * Find and fix the reliability problems that only appear past the point where upstream defaults and community guidance stop working, build each fix into the platform so the same class of problem does not come back, and automate the operational load that cannot be designed away. * Set the reliability standards for the platform: SLOs, error budgets, alerting philosophy, upgrade and rollout strategy, and the runbooks that follow from them. * Take on-call, lead incident response for the hardest production issues, and see the post-incident follow-up work through so the same incident does not happen twice. * Represent Apple in the upstream Kubernetes community, driving our requirements through design proposals and code in the relevant SIGs. * Grow the engineers around you through design review, code review, and direct mentoring. * Lead cross-team technical efforts through ambiguity, and advise partner service teams and their leadership on platform capabilities and tradeoffs, shaping both their designs and the Compute roadmap.

Minimum Qualifications

Bachelor's Degree in Computer Science, an engineering-related field, or equivalent related experience. Excellent verbal and written communication. You can write a design, iterate on it with the team, and present it to leadership in a way that clearly outlines the problem, solution, and impact. Deep experience building and operating core components of Kubernetes or a comparable orchestration system in production at scale: apiserver, etcd, scheduler, kubelet, controller-runtime, and the failure modes of the distributed Unix systems underneath them. Demonstrated ownership of a technical area end to end: you set the design, saw it through to production, and stayed accountable for how it behaved. Expert-level Golang, with a track record of shipping and owning controllers or operators that other teams depend on. Track record of raising the level of engineers around you through mentoring, design review, and the systems and tools you build.

Preferred Qualifications

Experience leading a technical effort spanning several teams with competing priorities and no clear owner at the start. Experience defining SLOs and error budgets that other teams adopted. Accepted contributions to upstream Kubernetes or a comparable open source project, or sustained participation in a SIG or working group. Deep familiarity with cloud-native observability such as Prometheus, Thanos, Grafana, or similar. Experience building AI-assisted tooling that other engineers adopted for triage, operational workflows, or infrastructure automation.