Valce Talent Solutions

 AWS Data Engineer

Valce Talent Solutions · Mexico

IT Services and IT Consulting · 11-50 employees

4 h ago
Remote Principal (10+ yrs) Full-time Mexico
Log in to apply, save this posting, or score it against your profile with AI.

About the role

Design, develop, and maintain automated ETL pipelines using Python and Apache Airflow. Optimize SQL queries for data warehouse performance and ensure data quality through rigorous testing and documentation.

What they look for

Python AWS SQL Snowflake Apache Airflow DBT ETL CI/CD GitHub Data Pipelines AWS S3 Data Warehousing Cloud Security Jinja Data Modeling

Requirements

Requires over 10 years of experience in data engineering with advanced proficiency in Python and SQL. Must have deep expertise in AWS cloud services and experience implementing CI/CD processes.

Full description

AWS Data Engineer

Remote

10+ Years Experience 

Great Communicator/Client Facing/ Attention to detail 

Individual Contributor and ability to work as a team. 

100% Hands on in the mentioned skills 

Programming Skills: 

Python and Data Pipelines: 

Advanced Proficiency in Python concepts like Code Structures, Modules, Packages, Class, SubClass, Inheritance, Multi-Threading and Functional Programming. 

Experience in developing reusable Python packages for internal or public usage 

Ability to write automating ETL processes and scheduling jobs like Airflow DAG. 

Ability to track job pipeline runs to reprocess error records 

Ability to orchestration different pipelines to run in sequence or parallel. 

Troubleshoot data pipeline errors and fix issues 

Export or Import data to/from various formats like CSV, JSON, XML etc preferably from S3 or other cloud storage. 

Experience in using AI IDE tool 

SQL: 

Advanced SQL skills, including complex joins, CTE's and subqueries 

Experience in optimizing SQL queries for performance and optimization in data warehouse technologies preferably Snowflake 

Testing and documentation: 

Proficiency in Python unit, integration and system test. 

Proficiency in implementing DBT tests for data validation and quality checks 

Code Generation: 

Experience in generating code using configurations using python and jinja templates 

Version control: 

Experience in GitHub, including implementing CI/CD process from scratch 

AWS Expertise: 

Data Storage solutions: 

In depth understanding of AWS S3 for data storage, ECS, IAM including best practices for organization and security 

Cloud Security: 

Knowledge of AWS security best practices, including IAM roles, encryption standard, secure coding guidelines, DBT profiles access configurations and more. 

Data Integration (nice to have): 

Experience with AWS lambda for serverless data processing tasks 

Workflow Orchestration (nice to have): 

Proficiency in using Apache Airflow on AWS to design ,schedule and monitor complex data flows 

Ability to integrate Airflow with AWS services and DBT models such as triggering a DBT model or EMR or reading from s3 writing to redshift or snowflake. Experience in API integrations to Salesforce, SFTP, SharePoint, One Drive and others. 

DBT Core/ Cloud Proficiency (nice to have): 

Experience in creating complex DBT models including full refresh, incremental models, snapshots, and documentation. Ability to write and maintain DBT macros for reusable code 

Experience in creating custom DBT macros using jinja and Python allowing for reusable components within DBT models 

Monitoring and Logging (nice to have): 

Familiarity with AWS cloud watch, Data Dog for monitoring the pipelines and setting up alerts for workflow failures