HEXAWARE

Big Data Engineer

HEXAWARE Mexico, Chihuahua, Mexico

IT Services and IT Consulting · 10,001+ employees

4 d ago
data-engineer Senior (5-10 yrs) Full-time Mexico
Log in to apply, save this posting, or score it against your profile with AI.

About the role

Build and maintain scalable ETL pipelines using Python and PySpark on AWS platforms. Collaborate with cross-functional teams to design storage solutions and implement data quality monitoring.

What they look for

Python PySpark AWS Glue SQL AWS Step Functions AWS Lambda Amazon Redshift ETL Pipelines Data Warehousing CI/CD Terraform GitLab Performance Engineering API Integration Data Quality CloudWatch

Requirements

Requires 4-8 years of software development experience with strong proficiency in Python, SQL, and AWS services. Candidates should have experience in data pipeline orchestration and performance optimization.

Full description

- Mid Level Dev AWS Data Engineer with 4-8 years of software development experience Build and maintain ETL pipelines using Python and PySpark on AWS Glue and related platforms. - Orchestrate workflows using AWS Step Functions and Lambda. - Implement messaging and event-driven integrations using SNS and SQS. - Design and optimize storage and querying solutions in Amazon Redshift, RDS, Oracle and S3-based architectures. - Write efficient SQL for transformations, validation, and reporting. - Integrate data from APIs and process structured and semi-structured JSON data. - Implement data quality checks, monitoring, and operational support processes. - Participate in CI/CD and version control practices for deployment and release management. - Collaborate with cross-functional teams to translate business requirements into technical solutions. - 4-8 years of software development experience across the appropriate platform. - Strong hands-on experience with Python, PySpark, API’s and SQL. - Experience with ETL/data pipeline development and Orchestration using Step functions / AirFlow . - Working knowledge of AWS services including Glue, Lambda, Step Functions, Redshift, S3, SNS, and SQS. - Experience with Athena, EMR, Kinesis, DynamoDB, or RDS. - Good Knowledge on CloudWatch, logging, and production support. - Understanding of data warehousing, data lakes, Lake House and query optimization. - Experience with GitLab/Terraform or similar and CI/CD workflows. - Good understanding of using AI tools like Github Copilot or similar for code productivity - Exposure to enterprise data lake or cloud migration initiatives. - Have an eye to solving complex problems, great communication with stakeholders - Have a good understanding of performance engineering of code pipelines and near real time systems - Good understanding on Agents and MCP"

Similar roles