Apple

Staff Data Engineer

Apple Austin, Texas, United States

Computers and Electronics Manufacturing · 10,001+ employees

4 h ago
data-engineer Principal (10+ yrs) Full-time United States
Log in to apply, save this posting, or score it against your profile with AI.

About the role

You will lead the data engineering and architecture for a centralized security data platform to manage risk and vulnerability findings. This involves owning data models, building production-grade pipelines, and creating services for ownership attribution and risk enrichment.

What they look for

Spark Java Python Scala Data Architecture Data Modeling SQL Distributed Systems AWS Hadoop Presto Flink Druid Data Engineering Security Risk Management Cloud Computing

Requirements

Candidates must have 15+ years of experience with distributed data technologies like Spark and proficiency in Java, Python, or Scala. Strong expertise in data architecture, modeling, and the ability to lead technical strategy within a complex, matrixed organization is required.

Full description

We are the Risk and Vulnerability Management (RVM) team in Apple Services Engineering (ASE) Security. We manage security risk for the infrastructure, platforms, and services behind iCloud, App Store, Apple Music, TV+, and Commerce. Findings reach us from scanners, red team engagements, design reviews, vendor advisories, threat intelligence, and bug bounty. Our job is to turn all of that into one prioritized backlog engineering teams can work from, and a posture picture leadership can trust.

Most of that job is data work. Signal arrives from dozens of systems at different quality and age, and it rarely says which asset, which service, or who owns it, so someone pieces that together by hand before anyone can act on it. That manual step sets the ceiling on how fast we identify and triage risk, and on how fast our partners can fix it. We are consolidating this onto one security data platform: a system of record for findings, a graph for ownership and blast radius, and a lakehouse for posture metrics and analytics.

Description

We are looking for a Staff Data Engineer to lead the data engineering and data architecture behind it. You will own the data model and the pipeline contracts other teams build against, move pipelines from prototype into production the business depends on, and make ownership attribution and risk enrichment a service instead of repeated one-off work. Success here takes deep distributed data engineering experience, real data judgment, and the ability to hold a technical direction across a large, matrixed engineering organization.

Minimum Qualifications

15+ years of experience working with Spark and other distributed data technologies (e.g. Hadoop, Presto, Flink, Druid) for building efficient & large scale data pipelines Highly proficient in at least one of Java, Python or Scala Deep expertise in Data Principles, Data Architecture & Data Modeling, Strong SQL skills Strong problem solver with meticulous attention to detail, capable of taking on loosely defined problems Experience working in a complex, matrixed organization involving cross-functional, and/or cross-business projects Strong communication and collaboration skills & ability to lead high-level discussions on technology strategy and approach Conceptually familiar with AWS cloud resources (S3, EC2, RDS etc) Conceptually familiar with OSI model and understanding of how network works (Load Balancer, Routers, network tagging, NetFlow) Conceptually familiar of the full technology stack (from BMCs, Firmware, to OS layer, to containers and applications) primitives

Preferred Qualifications

Experience with Cloud Computing platforms like Amazon AWS, Google Cloud Experience with building stream-processing applications using Apache Flink, Spark-Streaming, Apache Storm, Kafka Streams or others Experience with Search systems (such as ElasticSearch, Solr), NoSQL datastores (such as HBase, Cassandra, MongoDB) Experience building distributed, high-volume data services is a plus

Similar roles