Site Reliability Engineer
Apple Shanghai, Shanghai, China
Computers and Electronics Manufacturing · 10,001+ employees
About the role
You will manage critical integrations with supply chain partners and build robust, scalable DevOps platforms. The role involves leading incident response, performing root cause analysis, and driving reliability improvements across distributed systems.
What they look for
Requirements
Candidates must have at least 12 years of experience in SRE or DevOps with strong proficiency in programming languages like Java and Python. Deep expertise in Kubernetes, infrastructure as code, and database management is required to support large-scale production environments.
Benefits
Full description
Do you want to help build some of the largest and most consequential enterprise and customer technology systems in the world? Join Apple’s Information Systems and Technology (IS&T) organization. IS&T is the engine behind everything Apple does for customers and for the people who build for them. It’s Apple’s central nervous system. Supporting 2.5 billion active Apple devices, processing billions of secure transactions, and keeping the technology that defines modern life running flawlessly, IS&T makes the impossible feel effortless.
Do you love building solutions to handle global complexity and immense scale? Imagine what you could do here.
Sales and Operations Engineering is part of IS&T and drives the technology behind Apple's global operations, sales, and supply chain. The team connects the systems that move products from factory to customer — bridging sales platforms with the operational infrastructure that keeps Apple running at scale and ensuring Apple's technology and business priorities move in lockstep.
Description
Apple's B2B team manages critical integrations with Apple's supply chain partners such as manufacturers, logistics providers, banks, resellers and business customers. We are seeking a technically hands on individual with a real passion for programming and automation.
Join our dynamic team as a Software Reliability Engineer (SRE) and dive into innovative work culture fueled by machine learning, anomaly detection and threat detection. Collaborate with a highly motivated team of professionals who push boundaries and delivering exceptional results. This position offers an exciting opportunity to build your career as an SRE in a supportive environment, where continuous learning and professional development are prioritized.
Minimum Qualifications
At least 12 years of prior demonstrated experience in a Site Reliability Engineering, DevOps(Must), or an Infrastructure-focused role. Designing and Building DevOps platforms end-to-end alongwith SRE/Platform Engineering. Proven experience in building DevOps platforms from scratch. Applied Experience on GitOps-based deployment models (ArgoCD / Flux) Establish Infrastructure as Code (IaC) practices. Build and operate Kubernetes platforms (EKS / AKS / GKE / OpenShift) Experience working in large-scale, distributed systems Strong problem-solving and architectural skills Proficiency in one or more programming languages (eg. Python) Support of internet-facing production services and distributed systems via deployments, onCall and Incident Management. Lead incident response, RCA, and reliability improvements. Proficiency in implementing and coordinating telemetry using monitoring and observability tools like Splunk, Grafana, and Prometheus, or similar. Experience in solving and resolving issues in Kubernetes from both an operating system and application perspective. Building and operating container orchestrating systems like Kubernetes or EKS. Strong programming experience in Java building web, middleware or backend applications. Deep understanding of Oracle or similar relational databases and NoSQL databases such as MongoDB. Firsthand experience in performance tuning of applications and databases. Knowledge of HTTP/S, TCP, DNS, web application load balancing. Deep understanding of basic security concepts and protocols - authentication, authorization, signing, encryption, SSL/TLS, SSH/SFTP, PKI, X509 certificates and PGP.
Preferred Qualifications
Strong programming experience in Java for backend, middleware, or web applications Experience with NoSQL databases (MongoDB, Cassandra, DynamoDB, etc.) Deep understanding of relational databases (Oracle, PostgreSQL, MySQL, etc.) Hands-on experience in performance tuning of applications and databases Experience with advanced observability practices: 1) Distributed tracing ;2) SLO/SLI design; 3)Error budgets Prior experience in large-scale, highly distributed production environments Experience with container orchestration internals (scheduler, CNI, CSI, etc.) Knowledge of middleware platforms such as WebMethods Integration Server or similar. Experience with multi-cloud or hybrid cloud environments Familiarity with service mesh technologies (Istio, Linkerd, etc.)
Similar roles
-
SRE Engineering Manager
GitGuardian Paris, Ile-de-France, France
-
Site Reliability Engineer
NielsenIQ Mumbai, Maharashtra, India
-
Senior Site Reliability Engineer, Traffic Routing (Limited-time Relocation Bonus)
Vinted Kaunas, Kaunas County, Lithuania · €61K–€110K/yr
-
Head of DevOps & Site Reliability Engineering
Finastra pune, Maharashtra, India
-
Manager, Solution Engineering(Site Reliability Engineering)
Western Union pune, Maharashtra, India
-
Senior Site Reliability Engineer
Mastercard pune, Maharashtra, India