Apple

Site Reliability Engineer, Ai & Data Platforms

Apple · Shanghai, Shanghai, China

Computers and Electronics Manufacturing · 10,001+ employees

7 h ago
Mid (2-5 yrs) Full-time China
Log in to apply, save this posting, or score it against your profile with AI.

About the role

The SRE will operate and support production ETL pipelines, ensuring reliability, availability, and performance for data ingestion and transformation workflows. They will also triage production incidents, perform root cause analysis, and implement automation to improve system stability and reduce toil.

What they look for

Site Reliability Engineering Kubernetes Apache Spark Apache Airflow Kafka Snowflake Linux Python Bash Data Pipelines ETL CloudWatch Prometheus Grafana Incident Management SQL

Requirements

Candidates must have 4+ years of experience in SRE, DevOps, or platform engineering with strong hands-on skills in Linux, Kubernetes, Spark, and Airflow. Proficiency in Mandarin and English is required, along with the ability to troubleshoot distributed systems and manage production incidents.

Full description

We are hiring a Site Reliability Engineer to Build, support and improve Keystone, an ETL application/platform operating in the China region. Keystone supports data ingestion, transformation, and loading workflows across Kubernetes-based environments, including Airflow-based loader jobs that load data into Snowflake and Spark-on-EKS loader jobs that load data into a Datalake or Lakehouse environment.

The SRE will own reliability, availability, performance, data freshness, and operational readiness for production ETL pipelines. This role requires strong hands-on experience with Linux, Kubernetes, Spark, Airflow, Kafka or streaming systems, Snowflake, object storage, observability, and incident response.

The ideal candidate can troubleshoot distributed data platform issues from logs, metrics, timestamps, pipeline state, and infrastructure signals, and can drive permanent fixes through configuration improvements, automation, tuning, and engineering partnership.

This position is based in Shanghai, China.

Description

Key Responsibility

  • Operate and support production ETL pipelines, including extractors, loaders, batch jobs, streaming ingestion, and data consolidation workflows.
  • Support loader jobs running in an EKS Airflow cluster that load data into Snowflake.
  • Support Spark jobs running on EKS that load data into the Datalake or Lakehouse.
  • Own platform and pipeline reliability, including uptime, throughput, job success rate, SLA adherence, and data freshness.
  • Triage and resolve production incidents such as failed jobs, delayed loads, OOMKilled pods, CrashLoopBackOff pods, stalled pipelines, driver/executor failures, credential issues, and network/connectivity problems.
  • Perform root cause analysis for incidents and drive corrective actions through automation, configuration changes, capacity tuning, or code/platform improvements.
  • Tune Kubernetes workloads, including CPU/memory requests and limits, autoscaling, pod scheduling, service accounts, secrets, config maps, and workload health.
  • Tune Spark workloads, including executor count, cores, memory, memory overhead, shuffle partitions, dynamic allocation, adaptive query execution, spill reduction, and failed stage analysis.
  • Monitor and troubleshoot Airflow DAGs, task retries, scheduling issues, dependency failures, executor/resource constraints, and SLA misses.
  • Diagnose data latency across the full pipeline path, such as source event, Kafka or messaging layer, staging/object storage, ETL processing, Snowflake, and Datalake availability.
  • Support Kafka or equivalent streaming systems, including consumer lag, offsets, partitions, topic health, producer delay, and SASL/TLS client configuration.
  • Manage platform and job configuration through approved source-of-truth or GitOps-style processes instead of ad hoc production changes.
  • Manage secrets, certificates, truststores, JDBC credentials, Snowflake credentials, object-storage credentials, and key rotations following least-privilege practices.
  • Build and maintain dashboards, alerts, log queries, and runbooks for job health, freshness, backlog, failures, infrastructure usage, and incident response.
  • Write scripts and automation using Python, Bash, or similar tools to reduce toil and improve recovery speed.
  • Coordinate with application engineering, data engineering, platform, security, network, Snowflake, and Datalake teams.
  • Participate in on-call or production support rotation for China business hours and critical incidents.

Minimum Qualifications

4+ years of experience in SRE, DevOps, platform engineering, data infrastructure, or production operations. Strong Linux troubleshooting skills. Strong Kubernetes/EKS operations experience, including kubectl, deployments, statefulsets, pods, resource limits, service accounts, secrets, logs, events, and workload debugging. Hands-on experience supporting Apache Spark in production, preferably Spark on Kubernetes. Experience tuning Spark jobs, reading driver/executor logs, troubleshooting memory issues, shuffle failures, failed stages, and performance bottlenecks. Production experience with Apache Airflow, including DAG operations, task failures, retries, scheduling, SLA misses, and dependency troubleshooting. Practical experience with Kafka or similar messaging/streaming platforms, including consumer groups, lag, offsets, partitions, and secure client connectivity. Familiarity with Snowflake loading patterns, connectivity, roles, warehouses, stages, query/load history, and load monitoring. Familiarity with Datalake or Lakehouse architectures using object storage such as S3 or equivalent. Solid SQL skills and ability to analyze large operational or data pipeline datasets. Experience with monitoring and logging tools such as Splunk, Prometheus, Grafana, CloudWatch, ELK, Datadog, or equivalent. Scripting experience with Python and Bash. Strong incident management, RCA, change management, and operational documentation skills. Ability to troubleshoot distributed systems using evidence from logs, metrics, events, timestamps, and configuration. Professional working proficiency in Mandarin and English.

Preferred Qualifications

Experience supporting metadata-driven ETL platforms or internal ETL frameworks similar to Keystone. Experience with Iceberg, Hive Metastore, Parquet, ORC, or similar Lakehouse technologies. Experience with GitOps or source-of-truth configuration management. Experience with Snowflake performance tuning, warehouse sizing, query history analysis, and load optimization. Experience with Spark performance tuning at scale. Experience operating multi-region or region-specific data platforms. Experience with certificate, PKI, TLS, JKS/truststore, Kerberos, or key rotation processes. Experience with infrastructure as code tools such as Terraform, Helm, Argo CD, Ansible, or similar. Familiarity with API gateways, RabbitMQ, Redis, Cassandra, or other platform dependencies.