GSSTech Group

Senior Data Engineer - PySpark & Python

GSSTech Group · Bengaluru, Karnataka, India

IT Services and IT Consulting · 201-500 employees

8 h ago
Senior (5-10 yrs) Full-time India
Log in to apply, save this posting, or score it against your profile with AI.

About the role

Design, develop, and maintain scalable ETL pipelines and data marts using PySpark and Python for enterprise-scale analytics. Manage the full SDLC including development, UAT, production deployment, and performance optimization of complex SQL queries.

What they look for

Python PySpark ETL Pipeline Development Data Mart Development Data Warehousing Apache Spark SQL Apache Airflow CI/CD Git Data Analysis Feature Engineering NoSQL Hadoop Hive Pandas

Requirements

Requires over 5 years of commercial experience in data engineering with expert-level proficiency in Python and PySpark. Experience in banking or financial services and familiarity with Big Data technologies and software engineering best practices are highly preferred.

Full description

We are looking for an experienced Senior Data Engineer with strong expertise in PySpark and Python to join our Data Engineering team supporting enterprise-scale Data & Analytics initiatives. The ideal candidate will have hands-on experience in building scalable ETL pipelines, data marts, and production-grade data engineering solutions across structured, semi-structured, and unstructured datasets.

The role requires strong technical capabilities in Big Data technologies, data warehousing, data analysis, software engineering best practices, and end-to-end SDLC ownership. Candidates with banking or financial services domain experience will be highly preferred.

Key Responsibilities• Design, develop, and maintain scalable ETL pipelines and data marts using PySpark and Python.

  • Build robust, maintainable, and production-ready data engineering solutions.
  • Perform end-to-end SDLC activities including development, UAT support, bug fixes, production deployments, and post-production support.
  • Work with large-scale structured, semi-structured, and unstructured datasets.
  • Perform data analysis, cleansing, transformation, and feature engineering activities.
  • Debug and optimize PySpark code and complex SQL queries for performance and scalability.
  • Develop and maintain production-grade data pipelines using modern data engineering best practices.
  • Collaborate with cross-functional teams to resolve dependencies and ensure timely project delivery.
  • Participate in CI/CD implementation, testing, validation, and deployment activities.
  • Ensure data quality, integrity, and consistency across enterprise data platforms.
  • Work closely with technical and business stakeholders to understand data requirements and deliver scalable solutions.
  • Contribute to technical documentation, engineering standards, and process improvements.

Required Technical SkillsProgramming & Data Engineering• Python (Expert level)

  • PySpark (Expert level)
  • ETL Pipeline Development
  • Data Mart Development
  • Data Warehousing Concepts
  • End-to-End SDLC Experience

Big Data Technologies• Apache Spark (PySpark)

  • Hadoop
  • MapReduce
  • Hive
  • Pandas

Database Technologies• SQL

  • NoSQL Databases
  • Oracle SQL
  • Oracle Query Optimization & Data Analysis

Data Engineering & Analytics• Data Analysis

  • Data Cleansing
  • Data Linking
  • Data Transformation
  • Feature Engineering
  • Imputation Techniques
  • Data Validation

Workflow & Orchestration Tools• Apache Airflow

  • Oozie
  • Jenkins Pipelines

Software Engineering & DevOps• Git Version Control

  • CI/CD Pipelines
  • Testing & Validation of Data Pipelines
  • Production Deployment & Support
  • Software Engineering Best Practices

Development Tools• Jupyter Notebook

  • Git

Required Experience• 5+ years of commercial experience in Data Engineering or related data-driven roles.

  • Strong hands-on experience in building ETL pipelines and Data Marts.
  • Proven experience in developing production-grade PySpark and Python solutions.
  • Strong understanding of software engineering concepts and best practices.
  • Experience working with large-scale data processing frameworks.
  • Hands-on experience with production support, UAT activities, and deployment processes.
  • Strong analytical and debugging capabilities for PySpark and SQL-based data solutions.
  • Experience working within Agile delivery environments is preferred.

Preferred Domain Experience• Banking & Financial Services (Highly Preferred)

  • Digital Products
  • Data & Analytics Platforms

Soft Skills & Competencies• Strong analytical and problem-solving skills.

  • Excellent communication and interpersonal skills.
  • Ability to communicate effectively with both technical and non-technical stakeholders.
  • Strong ownership mindset and accountability for deliverables.
  • Ability to work under pressure and effectively prioritize tasks.
  • Strong collaboration skills with cross-functional teams.
  • Ability to lead technical initiatives and drive delivery outcomes.
  • Excellent verbal and written communication skills in English.

Nice to Have• Banking domain experience.

  • Experience working with enterprise-scale Data & Analytics platforms.
  • Exposure to Agile methodologies and modern data engineering practices.
  • Knowledge of production-grade data pipeline monitoring and optimization.