Apple

Data Engineer, IS&T Ai & Data Platforms

Apple Shanghai, Shanghai, China

Computers and Electronics Manufacturing · 10,001+ employees

8 h ago
data-engineer Senior (5-10 yrs) Full-time China
Log in to apply, save this posting, or score it against your profile with AI.

About the role

The role involves managing data migration and providing ongoing support for Applied Machine Learning enterprise systems in China. Responsibilities include setting up clusters, developing data pipelines, and enhancing applications for local data platform requirements.

What they look for

SQL Python Java Scala Apache Kafka Apache Spark Flink Tableau Streamlit Superset Snowflake PySpark Data Engineering Data Pipelines SRE Generative AI

Requirements

Candidates must have at least 5 years of data engineering experience and a bachelor's degree in a related field. Proficiency in Python, SQL, and distributed frameworks like Apache Spark and Kafka is required.

Full description

AI & Data Platforms (AiDP) is IS&T's engine for AI-powered innovation. The team brings together data, application development, and machine learning — including generative AI — along with data services and customer success functions, to help IS&T build solutions more efficiently and streamline the adoption and embedding of generative AI across Apple.

This position is required for working on the data migration and ongoing support for data platform for Applied Machine Learning enterprise systems at Apple China. The main deliverables will be:

  • Work with SRE on setting up and testing/validating new cluster, pipelines, applications, frameworks, connectivity and privacy and governance
  • Identify China specific datasets, migrate the datasets to China cluster
  • Modify and enhance existing applications and frameworks for China cluster

Description

The offered position requires at least 5 years subsequent data engineering experience including the following:

  • Writing advanced SQL for large scale relational data warehouses and data lakes.
  • Designing data pipelines for structured and semi-structured data.
  • Programing data engineering solutions in Python, Java, or Scala.
  • Working with distributed frameworks including Apache Kafka and Apache Spark and Flink
  • Designing reporting solutions with Tableau, Streamlit or Superset.

Minimum Qualifications

The offered position requires at least a bachelor's degree or equivalent in Computer Science, Information Systems, or a related field. The beneficiary has at least 5 years of subsequent data engineering experience across the required skills. Architecting scalable data processing systems for real-time, near-real-time, and batch data pipelines. Develop data engineering solutions in Python and advanced SQL. Develop self service data engineering applications for business users. Design complex Python, PySpark, SQL and data processing for large-scale data platforms including Snowflake and Apache Spark with SparkSQL.

Preferred Qualifications

Developing software system testing and validation procedures and documentation by writing efficient code that results in performant solutions; Conferring with data processing or project managers to obtain and refine system requirements. Monitoring existing software to correct errors, adapt it to new platform requirements, and implement new functionality.

Similar roles