AWS Data Engineer
Valce Talent Solutions · Mexico
IT Services and IT Consulting · 11-50 employees
About the role
Design, develop, and maintain automated ETL pipelines using Python and Apache Airflow. Optimize SQL queries for data warehouse performance and ensure data quality through rigorous testing and documentation.
What they look for
Requirements
Requires over 10 years of experience in data engineering with advanced proficiency in Python and SQL. Must have deep expertise in AWS cloud services and experience implementing CI/CD processes.
Full description
AWS Data Engineer
Remote
10+ Years Experience
Great Communicator/Client Facing/ Attention to detail
Individual Contributor and ability to work as a team.
100% Hands on in the mentioned skills
Programming Skills:
Python and Data Pipelines:
Advanced Proficiency in Python concepts like Code Structures, Modules, Packages, Class, SubClass, Inheritance, Multi-Threading and Functional Programming.
Experience in developing reusable Python packages for internal or public usage
Ability to write automating ETL processes and scheduling jobs like Airflow DAG.
Ability to track job pipeline runs to reprocess error records
Ability to orchestration different pipelines to run in sequence or parallel.
Troubleshoot data pipeline errors and fix issues
Export or Import data to/from various formats like CSV, JSON, XML etc preferably from S3 or other cloud storage.
Experience in using AI IDE tool
SQL:
Advanced SQL skills, including complex joins, CTE's and subqueries
Experience in optimizing SQL queries for performance and optimization in data warehouse technologies preferably Snowflake
Testing and documentation:
Proficiency in Python unit, integration and system test.
Proficiency in implementing DBT tests for data validation and quality checks
Code Generation:
Experience in generating code using configurations using python and jinja templates
Version control:
Experience in GitHub, including implementing CI/CD process from scratch
AWS Expertise:
Data Storage solutions:
In depth understanding of AWS S3 for data storage, ECS, IAM including best practices for organization and security
Cloud Security:
Knowledge of AWS security best practices, including IAM roles, encryption standard, secure coding guidelines, DBT profiles access configurations and more.
Data Integration (nice to have):
Experience with AWS lambda for serverless data processing tasks
Workflow Orchestration (nice to have):
Proficiency in using Apache Airflow on AWS to design ,schedule and monitor complex data flows
Ability to integrate Airflow with AWS services and DBT models such as triggering a DBT model or EMR or reading from s3 writing to redshift or snowflake. Experience in API integrations to Salesforce, SFTP, SharePoint, One Drive and others.
DBT Core/ Cloud Proficiency (nice to have):
Experience in creating complex DBT models including full refresh, incremental models, snapshots, and documentation. Ability to write and maintain DBT macros for reusable code
Experience in creating custom DBT macros using jinja and Python allowing for reusable components within DBT models
Monitoring and Logging (nice to have):
Familiarity with AWS cloud watch, Data Dog for monitoring the pipelines and setting up alerts for workflow failures