Senior Data Engineer - PySpark & Python
GSSTech Group · Bengaluru, Karnataka, India
IT Services and IT Consulting · 201-500 employees
About the role
Design, develop, and maintain scalable ETL pipelines and data marts using PySpark and Python for enterprise-scale analytics. Manage the full SDLC including development, UAT, production deployment, and performance optimization of complex SQL queries.
What they look for
Requirements
Requires over 5 years of commercial experience in data engineering with expert-level proficiency in Python and PySpark. Experience in banking or financial services and familiarity with Big Data technologies and software engineering best practices are highly preferred.
Full description
We are looking for an experienced Senior Data Engineer with strong expertise in PySpark and Python to join our Data Engineering team supporting enterprise-scale Data & Analytics initiatives. The ideal candidate will have hands-on experience in building scalable ETL pipelines, data marts, and production-grade data engineering solutions across structured, semi-structured, and unstructured datasets.
The role requires strong technical capabilities in Big Data technologies, data warehousing, data analysis, software engineering best practices, and end-to-end SDLC ownership. Candidates with banking or financial services domain experience will be highly preferred.
Key Responsibilities• Design, develop, and maintain scalable ETL pipelines and data marts using PySpark and Python.
- Build robust, maintainable, and production-ready data engineering solutions.
- Perform end-to-end SDLC activities including development, UAT support, bug fixes, production deployments, and post-production support.
- Work with large-scale structured, semi-structured, and unstructured datasets.
- Perform data analysis, cleansing, transformation, and feature engineering activities.
- Debug and optimize PySpark code and complex SQL queries for performance and scalability.
- Develop and maintain production-grade data pipelines using modern data engineering best practices.
- Collaborate with cross-functional teams to resolve dependencies and ensure timely project delivery.
- Participate in CI/CD implementation, testing, validation, and deployment activities.
- Ensure data quality, integrity, and consistency across enterprise data platforms.
- Work closely with technical and business stakeholders to understand data requirements and deliver scalable solutions.
- Contribute to technical documentation, engineering standards, and process improvements.
Required Technical SkillsProgramming & Data Engineering• Python (Expert level)
- PySpark (Expert level)
- ETL Pipeline Development
- Data Mart Development
- Data Warehousing Concepts
- End-to-End SDLC Experience
Big Data Technologies• Apache Spark (PySpark)
- Hadoop
- MapReduce
- Hive
- Pandas
Database Technologies• SQL
- NoSQL Databases
- Oracle SQL
- Oracle Query Optimization & Data Analysis
Data Engineering & Analytics• Data Analysis
- Data Cleansing
- Data Linking
- Data Transformation
- Feature Engineering
- Imputation Techniques
- Data Validation
Workflow & Orchestration Tools• Apache Airflow
- Oozie
- Jenkins Pipelines
Software Engineering & DevOps• Git Version Control
- CI/CD Pipelines
- Testing & Validation of Data Pipelines
- Production Deployment & Support
- Software Engineering Best Practices
Development Tools• Jupyter Notebook
- Git
Required Experience• 5+ years of commercial experience in Data Engineering or related data-driven roles.
- Strong hands-on experience in building ETL pipelines and Data Marts.
- Proven experience in developing production-grade PySpark and Python solutions.
- Strong understanding of software engineering concepts and best practices.
- Experience working with large-scale data processing frameworks.
- Hands-on experience with production support, UAT activities, and deployment processes.
- Strong analytical and debugging capabilities for PySpark and SQL-based data solutions.
- Experience working within Agile delivery environments is preferred.
Preferred Domain Experience• Banking & Financial Services (Highly Preferred)
- Digital Products
- Data & Analytics Platforms
Soft Skills & Competencies• Strong analytical and problem-solving skills.
- Excellent communication and interpersonal skills.
- Ability to communicate effectively with both technical and non-technical stakeholders.
- Strong ownership mindset and accountability for deliverables.
- Ability to work under pressure and effectively prioritize tasks.
- Strong collaboration skills with cross-functional teams.
- Ability to lead technical initiatives and drive delivery outcomes.
- Excellent verbal and written communication skills in English.
Nice to Have• Banking domain experience.
- Experience working with enterprise-scale Data & Analytics platforms.
- Exposure to Agile methodologies and modern data engineering practices.
- Knowledge of production-grade data pipeline monitoring and optimization.