About the role
You will own the full lifecycle of AWS-based data pipelines, from ingesting source data to delivering clean, consumer-ready datasets. Responsibilities include designing ETL/ELT workflows, managing S3 storage, and automating deployments using CI/CD pipelines.
What they look for
Requirements
Candidates must have 5 or more years of professional data engineering experience with a strong focus on AWS services like Glue, S3, and RDS. Proficiency in building data streaming pipelines with Kafka and managing CI/CD workflows via GitHub is essential.
Full description
About the Role
This is a fully hands-on, individual contributor Data Engineering role embedded within a data infrastructure initiative at a large enterprise. You will own the full lifecycle of AWS-based data pipelines, from ingesting and streaming source data through to delivering clean, consumer-ready datasets for downstream use.
What You'll Do
- Stream and process source data from mainframe and legacy systems, loading it into AWS S3.
- Design and implement ETL/ELT workflows using AWS Glue.
- Perform data reconciliation, validation, and quality checks to ensure data accuracy and reliability.
- Curate and transform data, provisioning clean datasets through AWS Aurora and RDS (PostgreSQL).
- Manage S3 storage including data retention policies, archival strategies, and lifecycle management.
- Automate workflows and manage deployments using GitHub and GitHub Actions-based CI/CD pipelines.
- Leverage AI tools to improve engineering productivity and automate data workflows.
What We're Looking For
- 5 or more years of professional Data Engineering experience building and delivering data pipelines, ETL/ELT workflows, or data platform solutions.
- Hands-on production experience with AWS Glue for designing and implementing ETL/ELT workflows.
- Strong working knowledge of AWS cloud services including S3, RDS/Aurora, and Glue.
- Demonstrated experience building and maintaining data streaming pipelines using Kafka.
- Experience implementing data reconciliation, data quality checks, and validation processes.
- Proficiency with GitHub repository management and GitHub Actions for CI/CD automation.
- Experience using AI tools to automate workflows and improve productivity.
- Ability to take ownership of deliverables and drive work to completion with minimal oversight.
- Experience with mainframe data sources or legacy system integration is a plus.
- Familiarity with MongoDB or other NoSQL databases is a plus.
Compensation & Benefits
This role pays $65/hr on a W2 basis.
Location
This role is 100% remote.
Similar roles
-
TS/SCI w/Poly - AI/ML Data Engineer
Leading Path Consulting Chantilly, Virginia, United States
-
Principal Consultant, Data Engineer
Lovelytics Chicago, Illinois, United States · $130K–$170K/yr
-
Database Administrator (Data Engineer)
Navy Federal Credit Union Pensacola, Florida, United States · $78K–$123K/yr
-
Data Engineer - AWS, GCP, Snowflake (CDI - H/F)
Talan Toulouse, Occitania, France
-
Data Engineer II
InVita Healthcare Technologies Baltimore, Maryland, United States · $100K–$120K/yr
-
Data Engineer
EXL Haryana, India