KPMG India

Associate Consultant - Data Engineer(Python/Pyspark)

KPMG India Bangalore, Karnataka, India

Business Consulting and Services · 10,001+ employees

4 h ago
data-engineer Mid (2-5 yrs) Full-time India
Log in to apply, save this posting, or score it against your profile with AI.

About the role

The candidate will design, develop, and maintain scalable data pipelines and architectures using Python, PySpark, and Azure services. They will collaborate with cross-functional teams to ensure data quality and support reporting and analytics initiatives.

What they look for

Python Pyspark SQL Azure Microsoft Fabric ETL/ELT Azure Data Factory Databricks Medallion Architecture Power BI Data Engineering Git Agile Delta Lake Azure Synapse Analytics

Requirements

Candidates must have strong experience in data engineering with Python, PySpark, and SQL, along with expertise in Azure cloud services. A degree in Computer Science or a related field is required, along with proficiency in ETL/ELT pipeline development.

Full description

We are seeking a highly skilled Data Engineer with expertise in Python, PySpark, SQL, Azure, and Microsoft Fabric. The successful candidate will be responsible for designing, developing, and maintaining scalable data pipelines and data architectures across cloud-based platforms.

Responsibilities:

  • Collaborate with cross-functional teams to design, develop, and implement scalable data solutions using Python, PySpark, SQL, and Azure services.
  • Design, develop, and maintain robust ETL/ELT pipelines for efficient data ingestion, transformation, and integration.
  • Leverage Azure Data Factory, Databricks, Microsoft Fabric to build and maintain modern data architectures such as medallion architecture.
  • Work closely with Power BI developers and data analysts to enable seamless data integration for reporting and analytics.
  • Participate in all phases of the Software Development Lifecycle (SDLC) including requirement gathering, development, testing, and deployment.
  • Write efficient, scalable, and optimized code to handle large volumes of structured and unstructured data.
  • Ensure data quality, consistency, and reliability across data pipelines and storage systems.
  • Collaborate with data scientists, analysts, and product teams to solve business and technical challenges.
  • Conduct code reviews and provide constructive feedback to team members.
  • Troubleshoot, debug, and resolve data pipeline and performance issues.
  • Stay up to date with industry trends and best practices in data engineering and cloud technologies.

Requirements:

  • Strong experience in data engineering using Python, PySpark, and SQL.
  • Proven expertise in designing and implementing ETL/ELT pipelines for data ingestion, transformation, and integration.
  • Solid understanding of relational databases with proficiency in SQL, including creating views and stored procedures.
  • Hands-on experience with Azure services such as Azure Data Factory, Azure Databricks, and Azure Synapse Analytics.
  • Experience implementing medallion architecture (Bronze, Silver, Gold layers) using Delta Lake in Azure Data Lake Storage (ADLS) or OneLake.
  • Familiarity with Microsoft Fabric for unified data integration, processing, and reporting.
  • Strong experience working with large-scale datasets and distributed computing frameworks like Spark.
  • Familiarity with version control systems such as Git.
  • Understanding of software development best practices and agile methodologies.
  • Strong problem-solving and analytical skills.
  • Excellent communication and collaboration skills.
  • Good to have knowledge of GenAI concepts including RAG architecture, AI agents, SKILL.md workflows, and AI orchestration frameworks.

Similar roles