Zensar

Data Engineer (AWS+Pyspark)

Zensar Hyderabad, Telangana, India

IT Services and IT Consulting · 10,001+ employees

Yesterday
data-engineer Mid (2-5 yrs) Full-time India
Create a free account to apply — email only, no card. You can also save this posting or score it against your profile with AI.

About the role

Build and optimize PySpark applications to process large volumes of data using AWS services. Manage data pipelines and ensure efficient storage and compression techniques are applied.

What they look for

Python Pyspark AWS Amazon EMR Amazon Athena AWS Glue Git Spark Dataframes Data warehousing Parquet Snappy Gzip Amazon Lambda Amazon EC2 Amazon S3 Amazon SNS

Requirements

Requires 4-5 years of experience in Big Data technologies with mandatory proficiency in Python and PySpark. Candidates should have hands-on experience with AWS analytics and compute services.

Full description

Data Engineer (AWS + pySpark)

  • Having 4-5 yrs years of relevant experience, which includes hands on experience in Big Data technologies.
  • Mandatory - Hands on experience in Python and PySpark.
  • Build pySpark applications using Spark Dataframes in Python.
  • Worked on optimizing spark jobs that processes huge volumes of data.
  • Hands on experience in version control tools like Git.
  • Worked on Amazon’s Analytics services like Amazon EMR, Amazon Athena, AWS Glue.
  • Worked on Amazon’s Compute services like Amazon Lambda, Amazon EC2 and Amazon’s Storage service like S3 and few other services like SNS.
  • Good to have knowledge of datawarehousing concepts – dimensions, facts, schemas- snowflake, star etc.
  • Have worked with columnar storage formats - Parquet etc. Well versed with compression techniques – Snappy, Gzip.
  • Good to have knowledge of AWS databases (atleast one) Aurora, RDS, Redshift, ElastiCache, DynamoDB.am

Similar roles