AWS Data Engineer
EXL Pune, Maharashtra, India
Business Consulting and Services · 10,001+ employees
About the role
Design and implement scalable, secure, and cost-optimized AWS data architectures while developing ETL pipelines using AWS Lambda and Glue. Orchestrate event-driven workflows and optimize data lakes using Apache Iceberg to support financial master data management.
What they look for
Requirements
Requires deep expertise in AWS cloud architecture, big data processing, and hands-on experience with Spark, PySpark, and Iceberg. Candidates must possess strong Python programming skills and a solid understanding of data lake optimization and federated querying.
Full description
We are seeking a highly skilled AWS Data Engineer with deep expertise in AWS cloud architecture, big data processing, real-time streaming, and modern data lake technologies. The ideal candidate will have strong hands-on experience in Spark (PySpark), Iceberg, EMR, Starburst/Trino, and event-driven architectures, along with experience building real-time and API-driven data applications who can design and build generic solutions for one of our Fortune 500 Client programs in the realm of Financial Master & Reference Data Management. This is high visibility, fast-paced key initiative will integrate data across internal and external sources, provide analytical insights, and integrate with the customer’s critical systems.
Responsibilities
Key Responsibilities
- Design and implement scalable, secure, and cost-optimized AWS data architectures.
- Develop and maintain ETL pipelines using AWS Lambda and AWS Glue ETL.
- Configure and manage AWS Glue Crawlers, Glue Data Catalog, and schema evolution.
- Build, optimize, and unit test applications on the Apache Spark framework using PySpark.
- Design and optimize data lakes using Apache Iceberg on AWS, including table compaction and Iceberg performance tuning.
- Work extensively with data formats such as Avro, Parquet, JSON, XML, and CSV.
- Orchestrate event-driven workflows using AWS Step Functions and Amazon EventBridge.
- Connect and integrate Starburst from Lambda and Glue ETL jobs for federated querying.
- Implement CI/CD pipelines for automated testing and deployment.
- Perform unit testing using PyTest, and performance tuning of Spark and Python applications
Qualifications
- Strong understanding of AWS architecture best practices, scalability, security, and cost optimization strategies.
- Strong hands-on experience with AWS services including Lambda, Glue ETL, Athena, S3, DynamoDB, Step Functions, EventBridge, SNS, and SQS.
- Deep experience in Apache Spark (PySpark/Scala) development, unit testing, and performance optimization.
- Strong Python programming skills using libraries such as pandas, requests, json, and awswrangler.
- Experience on Apache Kafka and Confluent Kafka.
- Experience designing and optimizing data lakes using Apache Iceberg, including compaction and Iceberg optimization techniques.
Similar roles
-
Senior AI-Native Data Engineer (f/m/x)
exmox Hamburg, Germany
-
Data Engineer
Ford Motor Company Chennai, Tamil Nadu, India
-
Data Engineer - DBT & SQL
Zensar Pune, Maharashtra, India
-
Data Engineer (REF5635Y)
Deutsche Telekom IT Solutions Budapest, Central Hungary, Hungary
-
Data Engineer
Entain Gibraltar, Gibraltar
-
Senior Data Engineer
Super Payments London, England, United Kingdom