Apple

Site Reliability Engineer / Devops — Retail Engineering

Apple Shanghai, Shanghai, China

Computers and Electronics Manufacturing · 10,001+ employees

4 h ago
sre Senior (5-10 yrs) Full-time China
Create a free account to apply — email only, no card. You can also save this posting or score it against your profile with AI.

About the role

You will drive the reliability, scalability, and deployment of compute platforms across hybrid cloud environments. Additionally, you will architect automation frameworks and lead testing strategies to ensure high availability and performance across retail systems.

What they look for

Site Reliability Engineering Java Python Go Infrastructure as Code Container Orchestration CI/CD Distributed Systems Observability Prometheus Grafana Datadog OpenTelemetry Kafka Kubernetes Cloud Computing

Requirements

Candidates must have a Bachelor's degree in Computer Science or equivalent with at least 7 years of experience in Site Reliability Engineering. Proficiency in modern programming languages and experience with large-scale distributed systems and observability stacks are required.

Full description

Apple’s IS&T Retail Engineering team is seeking a SDET to own quality across our retail ecosystem — spanning eCommerce backend services, SAP integrations, and cross-functional end-to-end workflows. This is a hands-on technical leadership role: you’ll architect automation frameworks, lead testing strategy across multiple concurrent projects, and drive quality standards organization-wide.

You’ll work at the intersection of backend development, test automation, and program coordination — ensuring that features shipping to apple.com/shop and supporting retail systems meet Apple’s bar for reliability, performance, and customer experience.

Description

You will drive the reliability, deployment, and scalability of compute platforms across on-premises and hybrid cloud environments. Collaborating closely with cross-functional technical and business partners, you will build Infrastructure as Code, optimize container orchestration, and streamline CI/CD delivery pipelines. You will champion automation and operational excellence, ensuring high availability, robust security standards, and proactive observability across large-scale distributed systems.

Minimum Qualifications

Bachelor’s degree in Computer Science or equivalent field with 7+ years of experience, or Master’s degree with 5+ years of experience. 7+ years of experience in Site Reliability Engineering with a strong focus on building, scaling, and operating large-scale distributed platform services, and Java applications. Strong technical grasp of Open Source technologies designed for large-scale data processing. Proven expertise in designing, analyzing, and troubleshooting complex distributed systems. Proficiency in at least one modern programming or scripting language (Python, Java, Go, Bash, Ansible, or similar). Practical experience designing and deploying end-to-end observability stacks (Prometheus, Grafana, Datadog, OpenTelemetry, ELK, etc.). Demonstrated troubleshooting and problem-solving skills across production software and infrastructure environments.

Preferred Qualifications

In-depth understanding of SRE principles, including error budgeting, SLO/SLI/SLA definition, and advanced observability practices (Prometheus, Splunk, Grafana, OpenTelemetry). Advanced programming skills in Java, Python, or Go, with hands-on experience across relational, NoSQL, or OLAP databases and event-driven streaming architectures (Kafka, RabbitMQ). Track record of managing production on-call rotations, critical incident triage, root cause analysis (RCA), and post-incident reviews (PIR). Solid knowledge of enterprise security standards, cryptography, authentication protocols (OAuth, SAML, SSO), and compliance governance.

Similar roles