Site Reliability Engineer — Data Platforms, IS&T Ai & Data Platforms
Apple Shanghai, Shanghai, China
Computers and Electronics Manufacturing · 10,001+ employees
About the role
You will operate and harden large-scale data and ML platforms while driving down MTTR and automating operational toil. You are responsible for performance tuning, capacity planning, and resolving production incidents to ensure high availability and reliability.
What they look for
Requirements
Candidates must have 3+ years of experience in large-scale distributed systems and proficiency in programming languages like Python, Go, or Java. A strong understanding of reliability principles, Kubernetes, and data processing frameworks like Spark or Flink is required.
Full description
The AiDP Data Platforms team builds and operates data platforms at scale on the Cloud, helping Apple process, store, and access petabytes of data. We’re seeking an SRE to own the reliability, performance, and operability of our data and ML platforms — someone who thinks in SLOs, failure modes, and blast radius, and is passionate about keeping large-scale distributed systems fast, available, and cost-efficient.
Description
You’ll operate and harden our big data platform — built on open source and other technologies — that powers critical applications like analytics, reporting, and AI/ML. This means driving down MTTR, automating operational toil, tuning performance and cost, running capacity planning, and root-causing production incidents before and after they happen. You’re an independent, self-directed problem-solver who communicates clearly with both engineers and non-technical partners, and you’ll work across many teams to keep the platform running at Apple’s standard.
Minimum Qualifications
3+ years of experience operating and supporting critical, large-scale distributed systems in production, with scripting/programming ability in Python, Go, Java, Scala, or Bash for automation and tooling. Deep understanding of reliability principles — fault tolerance, high availability, low latency, graceful degradation — and how to instrument and enforce them (SLIs/SLOs, alerting, on-call practices, incident response, postmortems). Hands-on experience operating data processing ecosystems and distributed computing frameworks (Spark, Flink) and MPP query engines (Trino, StarRocks), including performance tuning and capacity management. Proficiency operating Kubernetes/Helm at scale, building and maintaining CI/CD pipelines (GitHub Actions, Jenkins), managing infrastructure as code (Terraform, Pulumi), and running service-oriented architectures across multi-cloud environments. Strong troubleshooting and performance analysis skills in complex production environments; fluency in Unix/Linux and command-line diagnostics, with excellent problem-solving and communication skills.
Preferred Qualifications
Experience contributing to open source projects or operating across multiple public cloud providers. In-depth operational knowledge of specific distributed frameworks — Spark, Flink, or Kafka Streams, Trino, Iceberg — including debugging via component-specific logs, and experience managing multi-tenant Kubernetes clusters at scale. Experience with workflow/pipeline orchestration tools (Airflow, dbt) and understanding of data modeling and warehousing concepts. Experience debugging Kubernetes/Spark production issues via logs and metrics, with a continuous-improvement mindset for self, team, and org. Bachelor’s or Master’s degree in Computer Science, Computer Engineering, or a related field.
Similar roles
-
Senior Site Reliability Engineer
Pythian Argentina
-
SRE / Infrastructure Engineer (LABGEN)
Medfar Village of Great Neck Plaza, New York, United States · $90K–$115K/yr
-
SRE / Infrastructure Engineer (MEDGEN)
Medfar Village of Great Neck Plaza, New York, United States · $90K–$115K/yr
-
Lead Site Reliability Engineer
Mattermost United States
-
Site Reliability Engineer - Core
Blockchain.com Buenos Aires, Argentina
-
Site Reliability Engineer
Pythian Argentina