KPMG India

Associate Consultant - Data Engineer

KPMG India Gurgaon, Haryana, India

Business Consulting and Services · 10,001+ employees

7 h ago Closes today
data-engineer Mid (2-5 yrs) Full-time India
Log in to apply, save this posting, or score it against your profile with AI.

About the role

Design, develop, and maintain scalable data pipelines and engineering solutions using Databricks, PySpark, and SQL to support business analytics. Collaborate with cross-functional teams to translate business requirements into secure, cost-effective cloud data architectures.

What they look for

Databricks PySpark Python Advanced SQL Data Engineering ETL/ELT Data Modeling AWS Azure Microsoft Fabric Data Warehousing Git CI/CD Power BI DBT

Requirements

Requires over 2 years of hands-on experience in data engineering, data warehousing, and large-scale data processing. Candidates must hold a Bachelor's or Master's degree in a relevant technical field and possess strong expertise in cloud-based data platforms.

Full description

We are recruiting for an Associate Consultant in the team.

Your responsibilities will include:

  • Design, develop, and maintain scalable data pipelines and data engineering solutions using Databricks, PySpark, Python, and Advanced SQL to support business reporting and analytics needs.
  • Build and optimize ETL/ELT processes for ingesting, transforming, and processing structured and semi-structured data while ensuring data quality and reliability.
  • Develop and maintain robust data models and curated datasets aligned with business and analytical requirements.
  • Design and implement scalable, secure, and cost-effective data architectures on AWS and Azure, leveraging cloud-native services, Databricks, and Microsoft Fabric to support enterprise analytics and reporting requirements.
  • Work closely with business stakeholders to understand requirements, ask appropriate business questions, and translate them into effective technical solutions.
  • Perform data validation, testing, monitoring, and troubleshooting to ensure accuracy, performance, scalability, and data integrity across pipelines.
  • Leverage Databricks best practices to optimize Spark jobs, SQL workloads, and overall platform performance.
  • Collaborate with cross-functional teams including analysts, architects, and business users to deliver high-quality data solutions and support data-driven decision making.
  • Ensure all solutions adhere to organizational standards, security guidelines, governance policies, and development best practices

>> SKILLS & QUALIFICATIONS:

To succeed in this demanding role you will need to demonstrate the following skills and experience:

  • Over 2 years of hands-on experience in Data Engineering, DataWarehousing, and large-scale data processing projects.
  • Strong expertise in Databricks/ MS Fabric, Advanced SQL,Python, PySpark, Data Pipeline Development, and DataModeling. Power BI experience will be a plus.
  • Experience in designing and developing scalable ETL/ELTpipelines and working with structured, semi-structured, andunstructured data.
  • Strong understanding of data warehousing concepts, dimensional modeling, performance optimization, and modern data architectures.
  • Hands-on experience with cloud-based data platforms, preferably AWS and Azure, including cloud storage, compute, orchestration, and data integration services.
  • Experience with Git/version control, CI/CD practices, and cloud-based data platforms; exposure to Microsoft Fabric and DBT is preferred.
  • Ability to translate business requirements into scalable technical solutions with a focus on performance, reliability, and maintainability.

Qualification:

  • Bachelor's/Master's degree in Computer Science, InformationTechnology, Engineering, or a related field.
  • Relevant Databricks or Microsoft certifications are an addedadvantage.

Similar roles