Amazon

Senior Data Engineer, Infrastructure Reliability

Amazon · Austin, Texas, United States · $155K–$209K/yr

Software Development · 10,001+ employees

7 h ago
Senior (5-10 yrs) Full-time United States
Log in to apply, save this posting, or score it against your profile with AI.

About the role

You will design, build, and operate scalable ETL pipelines to ingest and transform telemetry data for ML models and incident detection. Additionally, you will collaborate with applied scientists to develop a unified data lake architecture and ensure the reliability of production data systems.

What they look for

Data engineering Data modeling ETL pipelines SQL Python Java Scala NodeJS Big data Hadoop Hive Spark EMR Data warehousing CloudWatch DynamoDB

Requirements

Candidates must have at least 5 years of data engineering experience, proficiency in SQL, and experience with at least one modern programming language. A bachelor's degree in a technical field such as computer science or engineering is required.

Benefits

Medical coverage Dental coverage Vision coverage Maternity leave Parental leave Paid time off 401(k) plan Life insurance Mental health support Flexible spending accounts Adoption reimbursement Surrogacy reimbursement Restricted stock units

Full description

Help build the data foundation that keeps Amazon's fulfillment network running 24/7. Infrastructure Reliability is building an AI-powered platform that detects, diagnoses, and resolves incidents across thousands of sites globally, and none of it works without clean, reliable, well-modeled data.

You will own and evolve production data pipelines that power real-time incident detection and correlation, build the data lake and unified data models that unblock ML and autonomous resolution initiatives, and rescue historical telemetry data before it expires and becomes permanently unrecoverable. This is a high-ownership role with direct, visible impact on Amazon's global fulfillment operations, working closely with applied scientists who train models on the data you build.

Key job responsibilities You will design, build, and operate ETL pipelines that ingest, transform, and correlate incident and telemetry data from sources including DynamoDB, OpenSearch, CloudWatch, and internal incident management systems, publishing curated datasets to our data lake and Andes for consumption by dashboards, applications, and ML models. You will take ownership of production pipelines, including scheduled batch processing jobs and Glue-based ETL workflows, ensuring reliability, monitoring, and timely resolution of pipeline issues.

You will build data retention and ingestion pipelines to preserve time-bound telemetry signals before source-system expiration windows, and you will design standardized ingestion frameworks that normalize data from many disparate sources into a common, science-consumable format. You will contribute to the design of a unified data lake architecture, including schema design, partitioning strategy, and access patterns, replacing fragmented, duplicated pipelines with a single source of truth.

You will partner closely with applied scientists and engineers to grasp data requirements for ML model training and feature engineering, and you will mentor other engineers on data engineering best practices, code quality, and pipeline design.

A day in the life You might start your day investigating a pipeline failure alert, tracing it back through Glue job logs to a schema change upstream. Later, you're in a design discussion with an applied scientist about what shape of data would best enable a new model, translating that into a concrete schema. In the afternoon, you're heads-down building a new ingestion adapter or reviewing a teammate's pull request. Your work directly determines what data is available, and reliable, for the platform's detection and reasoning capabilities.

Amazon offers a full range of benefits that support you and eligible family members, including domestic partners. Benefits can vary by location, the number of regularly scheduled hours you work, length of employment, and job status such as seasonal or temporary employment. The benefits that generally apply to regular, full-time employees include: 1. Medical, Dental, and Vision Coverage 2. Maternity and Parental Leave Options 3. Paid Time Off (PTO) 4. 401(k) Plan

If you are not sure that every qualification on the list above describes you exactly, we'd still love to hear from you! At Amazon, we value people with unique backgrounds, experiences, and skillsets. If you’re passionate about this role and want to make an impact on a global scale, please apply!

About the team Infrastructure Reliability sits within Amazon's Robotics organization, building the platform that keeps fulfillment operations running no matter what breaks. We do not own any single domain; we build the data and orchestration layer that sees across all of them, identifying failures that cascade across team boundaries. We are a small, technically deep team building AI-powered detection and remediation capabilities at scale. We value ownership, rigor, and hands-on technical depth, and we move quickly from idea to production.

Basic Qualifications: - 5+ years of data engineering experience - Experience with data modeling, warehousing and building ETL pipelines - Experience with SQL - Experience in at least one modern scripting or programming language, such as Python, Java, Scala, or NodeJS - Experience mentoring team members on best practices - Bachelor's degree in computer science, engineering, analytics, mathematics, statistics, IT or equivalent

Preferred Qualifications: - Experience with big data technologies such as: Hadoop, Hive, Spark, EMR - Experience operating large data warehouses - Master's degree in computer science, engineering, analytics, mathematics, statistics, IT or equivalent

Amazon is an equal opportunity employer and does not discriminate on the basis of protected veteran status, disability, or other legally protected status.

Our inclusive culture empowers Amazonians to deliver the best results for our customers. If you have a disability and need a workplace accommodation or adjustment during the application and hiring process, including support for the interview or onboarding process, please visit https://amazon.jobs/content/en/how-we-hire/accommodations for more information. If the country/region you’re applying in isn’t listed, please contact your Recruiting Partner.

The base salary range for this position is listed below. Your Amazon package will include sign-on payments and restricted stock units (RSUs). Final compensation will be determined based on factors including experience, qualifications, and location. Amazon also offers comprehensive benefits including health insurance (medical, dental, vision, prescription, Basic Life & AD&D insurance and option for Supplemental life plans, EAP, Mental Health Support, Medical Advice Line, Flexible Spending Accounts, Adoption and Surrogacy Reimbursement coverage), 401(k) matching, paid time off, and parental leave. Learn more about our benefits at https://amazon.jobs/en/benefits.

USA, TX, Austin - 154,600.00 - 209,100.00 USD annually