Apple

Site Reliability Engineer

Apple Shanghai, Shanghai, China

Computers and Electronics Manufacturing · 10,001+ employees

20 h ago
sre Mid (2-5 yrs) Full-time China
Log in to apply, save this posting, or score it against your profile with AI.

About the role

This role is responsible for the reliability, monitoring, and operational health of production data center services. It involves triaging incidents, building automation to reduce manual work, and partnering with engineering teams to improve system resilience.

What they look for

Site Reliability Engineering DevOps Automation Infrastructure Management Hybrid Infrastructure Security Practices Distributed Systems Troubleshooting Observability Container Orchestration CI/CD Pipelines Oracle Cassandra MongoDB Postgres Scripting

Requirements

Candidates must have a bachelor's degree in Computer Science or a related field and at least 3 years of experience in SRE, DevOps, or production operations. Proficiency in English and Mandarin is required, along with experience in infrastructure automation and security practices.

Full description

Joint Mobile Engineering Team (JMET) is a security engineering team that provides critical services for Apple across every product line. From manufacturing to customer-facing operations, the team's services span the entire lifecycle of most Apple hardware. The team designs, implements, and supports services that improve customer safety and privacy through security services tightly coupled with hardware — including server-side solutions that activate Apple devices worldwide and support Apple's efforts in eSIM. The team works closely with cross-functional teams across Apple, as well as carriers and other third parties. Many of the team's services are referenced in the iOS Security Guide or discussed publicly online.

As a Site Reliability Engineer, you will participate in initiatives that are important to the success of upcoming product launches and security initiatives. You will also work with large cross-functional teams to align expectations and validate the work you're doing.

Description

This role is responsible for the reliability, monitoring, and operational health of production data center services. It includes triaging and resolving incidents, building automation to reduce manual work, and partnering with engineering teams to improve system resilience as services scale.

Minimum Qualifications

Bachelor's degree in Computer Science, a related field, or equivalent practical experience 3+ years of experience in a Site Reliability Engineering, DevOps, or production operations role Experience automating infrastructure tasks using one or more programming or scripting languages Experience with hybrid infrastructure management across on-prem and cloud environments (e.g., AliCloud): compute, networking, storage, and data stores such as Oracle, Cassandra, MongoDB, and Postgres Experience applying security practices: access controls, patching, TLS/SSL, and identifying threats at the application and network level Experience troubleshooting distributed systems failures and driving incidents through to resolution Proficiency in English and Mandarin

Preferred Qualifications

Experience with observability and alerting tools (e.g., metrics, tracing, and dashboarding platforms) Experience applying SRE principles — error budgets, SLAs, SLOs, and SLIs — to measure and improve service reliability Experience with container orchestration platforms, including deployments, networking, storage, and secrets management Foundational knowledge of release engineering: CI/CD pipelines, version control, and deployment methodologies Foundational knowledge of distributed systems concepts: scalability, fault tolerance, and trade-offs under load Ability to communicate clearly across cross-functional and global teams, adjusting technical detail for different audiences Ability to work independently and take initiative in ambiguous or fast-changing situations

Similar roles