SRE Engineer
Apple Shanghai, Shanghai, China
Computers and Electronics Manufacturing · 10,001+ employees
About the role
You will design, build, and scale a modern SRE ecosystem while implementing best-in-class DevOps practices for Apple's B2B systems. Responsibilities include managing machine learning model lifecycles, improving system reliability, and leading incident response and performance optimization efforts.
What they look for
Requirements
The role requires deep hands-on expertise in building DevOps platforms from scratch and strong architectural thinking in cloud-native environments. Candidates must possess proficiency in programming languages like Python and Java, along with extensive experience in Kubernetes, IaC, and observability tools.
Full description
Do you want to help build some of the largest and most consequential enterprise and customer technology systems in the world? Join Apple’s Information Systems and Technology (IS&T) organization. IS&T is the engine behind everything Apple does for customers and for the people who build for them. It’s Apple’s central nervous system. Supporting 2.5 billion active Apple devices, processing billions of secure transactions, and keeping the technology that defines modern life running flawlessly, IS&T makes the impossible feel effortless.
Do you love building solutions to handle global complexity and immense scale? Imagine what you could do here.
Customer Systems is part of IS&T and drives the technology behind Apple's customer support experience — from contact center operations to the software powering the iconic Genius Bar. The team also builds and operates AppleCare's online support platform, which handles 6 billion visits per year, delivering seamless, high-quality support to Apple customers around the globe.
Description
As an SRE at Apple, you will be part of a team who will implement and maintain best-in-class devops practices, work on complex technical challenges related to scalability, reliability and performance of Apple B2B systems. You will be managing the lifecycle of machine learning models in production and non-production environment. You will be responsible for continuously assessing and improving system processes, detecting anomalies, identify the areas of optimization and implementing solutions to enhance system reliability and performance. You should have a passion for programming and a good conceptual understanding of the operating environment - JVM, Operating System, File Systems, Network Protocols. Technical expertise, strong communication skills and teamwork are essential requirements for this role as it involves working with both technical and non-technical groups within Apple and externally with our supply chain partners.
We are looking for a Senior SRE Engineer who can design, build, and scale a modern SRE ecosystem from scratch. This role requires deep hands-on expertise, strong architectural thinking, and the ability to establish GitOps-driven, cloud-native CI/CD platforms using the latest technologies. The ideal candidate will act as a foundational engineer and technical leader, defining standards, tooling, automation, and reliability practices across the organization.
Minimum Qualifications
Designing and Building DevOps platforms end-to-end along with SRE/Platform Engineering. Proven experience in building DevOps platforms from scratch. Applied Experience on GitOps-based deployment models (ArgoCD / Flux) Establish Infrastructure as Code (IaC) practices. Build and operate Kubernetes platforms (EKS / AKS / GKE / OpenShift) Experience working in large-scale, distributed systems Strong problem-solving and architectural skills Proficiency in one or more programming languages (eg. Python) Support of internet-facing production services and distributed systems via deployments, onCall and Incident Management. Lead incident response, RCA, and reliability improvements. Proficiency in implementing and coordinating telemetry using monitoring and observability tools like Splunk, Grafana, and Prometheus, or similar. Experience in solving and resolving issues in Kubernetes from both an operating system and application perspective. Building and operating container orchestrating systems like Kubernetes or EKS. Strong programming experience in Java building web, middleware or backend applications. Deep understanding of Oracle or similar relational databases and NoSQL databases such as MongoDB. Firsthand experience in performance tuning of applications and databases. Knowledge of HTTP/S, TCP, DNS, web application load balancing. Deep understanding of basic security concepts and protocols - authentication, authorization, signing, encryption, SSL/TLS, SSH/SFTP, PKI, X509 certificates and PGP.
Preferred Qualifications
Strong programming experience in Java for backend, middleware, or web applications Experience with NoSQL databases (MongoDB, Cassandra, DynamoDB, etc.) Deep understanding of relational databases (Oracle, PostgreSQL, MySQL, etc.) Hands-on experience in performance tuning of applications and databases Experience with advanced observability practices: -Distributed tracing -SLO/SLI design -Error budgets Prior experience in large-scale, highly distributed production environments Experience with container orchestration internals (scheduler, CNI, CSI, etc.) Knowledge of middleware platforms such as WebMethods Integration Server or similar. Experience with multi-cloud or hybrid cloud environments Familiarity with service mesh technologies (Istio, Linkerd, etc.)
Similar roles
-
Site Reliability Engineer | Weekend Warrior
Jump Trading Amsterdam, North Holland, Netherlands · €150K–€175K/yr
-
[MLA] Senior Site Reliability Engineer (SRE) – Kubernetes
Software Mind Krakow, Lesser Poland Voivodeship, Poland
-
Senior Site Reliability Engineer
Mozn Cairo, Cairo, Egypt
-
Manager- Site Reliability Engineering
Okta Bengaluru, Karnataka, India
-
Senior Site Reliability Engineer (12m FTC)
Mantel Sydney, New South Wales, Australia
-
Senior Site Reliability Engineer
2K Bangalore, Karnataka, India