Senior Data Engineer
Friends Color Images Pvt Ltd Noida, Uttar Pradesh, India
Consumer Services · 201-500 employees
About the role
The Senior Data Engineer will design and manage scalable data pipelines to transform complex banking transaction data into production-ready datasets. They will also collaborate with ML and engineering teams to ensure data quality, performance optimization, and secure deployment within restricted banking environments.
What they look for
Requirements
Candidates should have 5-9 years of experience in data engineering with strong expertise in Python, SQL, and distributed processing frameworks like Spark. Proficiency in orchestration tools, data modeling, and production-grade pipeline management is essential for this role.
Benefits
Full description
Broad Function:
The Senior Data Engineer – VARTA SENSE will be responsible for designing, developing, and managing scalable and reliable data pipelines that form the core data foundation of the VARTA SENSE platform.
The role will work closely with the ML Lead, Solution Architect, Backend Engineering team and other technology stakeholders to transform complex banking and transaction data into standardized, reusable and production-ready datasets and features.
The position requires strong hands-on expertise in Python, SQL, Spark/PySpark, data pipelines, orchestration, data quality and distributed data processing, with particular emphasis on reliability, performance, scalability and secure deployment within restricted banking environments.
Roles and Responsibilities (not limited to):
1. Data Pipeline Development & Engineering
- Design and develop scalable batch, incremental and production-grade data pipelines for large volumes of transaction and customer data.
- Build standardized and reusable data models and canonical schemas for transactions, customer attributes, offers, exposure and outcome events.
- Develop configurable ingestion and mapping frameworks to integrate data received from different banks and source systems.
- Manage pipeline orchestration including dependencies, checkpointing, retries, partial failures, backfills and safe replay mechanisms.
- Implement deduplication, late-arriving data handling, incremental loads, CDC, watermarking and merge/upsert processing.
2. Data Quality, Reconciliation & Governance
- Establish strong data-quality controls, automated testing, source-to-target reconciliation and schema validation.
- Ensure data lineage and schema evolution are properly managed across pipelines.
- Proactively identify and prevent issues such as missing data, duplicate records, incorrect aggregations, silent data loss and double counting.
- Maintain reliable and auditable data-processing standards across production environments.
3. ML Feature Engineering Support
- Work closely with the ML Lead to develop and maintain analytical and ML feature pipelines.
- Ensure consistency between training and production/inference datasets.
- Maintain point-in-time correctness and prevent data leakage within feature pipelines.
- Support versioned analytical features and reusable feature definitions.
4. Performance & Scalability
- Optimize large-scale data-processing workloads through effective partitioning, query optimization, join strategies, memory management and I/O optimization.
- Benchmark workloads against expected transaction volumes and available infrastructure.
- Identify bottlenecks and continuously improve processing speed, stability and infrastructure utilization.
- Ensure data pipelines remain scalable as transaction volumes and customer deployments increase.
5. Deployment, Security & Production Reliability
- Develop and support pipelines for deployment within bank-controlled on-premise and private-cloud environments.
- Ensure pipelines can operate in restricted or offline environments without public-internet dependency at runtime.
- Implement appropriate controls for PII handling, encryption, tokenization, masking, access management and secure connectivity.
- Establish monitoring and alerting for pipeline failures, data-quality issues, infrastructure bottlenecks and processing delays.
- Define and maintain backfill, replay, recovery and disaster-recovery procedures for critical pipelines.
6. Cross-functional Collaboration
- Work closely with ML, Architecture, Backend Engineering, DevOps and Infrastructure teams to design and deliver reliable data solutions.
- Collaborate with development partners to reproduce, transition and operationalize data pipelines within FCI.
- Support integration between data platforms and downstream application / ML services.
- Participate in technical discussions, architecture reviews and production troubleshooting.
7. Reusability, Documentation & Knowledge Transfer
- Build reusable adapters and configuration layers so that onboarding a new bank can be managed through configuration and mapping rather than changes to the core product code.
- Maintain clear technical documentation, runbooks and operational procedures.
- Document architecture decisions, troubleshooting steps and recovery procedures.
- Enable alternate engineering resources to independently operate and troubleshoot critical pipelines.
- Participate in code reviews, Git-based development, CI/CD practices and continuous improvement of engineering standards.
Requirements
Desired Qualifications & Experience:
- 5–9 years of relevant experience in Data Engineering, with hands-on ownership of production-grade data pipelines.
- Strong hands-on expertise in Python and Advanced SQL.
- Strong understanding of SQL concepts including window functions, query optimization and query-plan analysis.
- Practical experience with distributed data-processing platforms such as Apache Spark / PySpark or equivalent technologies.
- Experience building and managing production pipelines using Apache Airflow or an equivalent orchestration framework.
- Strong understanding of data modelling, ETL/ELT, batch processing and incremental data pipelines.
- Experience working with Parquet and other columnar data formats, partitioning, schema evolution and reconciliation.
- Practical experience with CDC / incremental ingestion technologies and concepts.
- Exposure to technologies such as Kafka, Debezium or equivalent event / CDC platforms would be preferred.
- Experience implementing automated Data Quality frameworks using tools such as Great Expectations, dbt tests, Soda or equivalent.
- Understanding of Data Lineage and Metadata Management using OpenLineage, Marquez or similar self-hosted solutions would be an advantage.
- Experience with lakehouse technologies such as Delta Lake, Apache Iceberg or Apache Hudi would be preferred.
- Exposure to Feature Store platforms such as Feast or equivalent would be beneficial.
- Working knowledge of Git, CI/CD, Docker and Linux environments.
- Experience with monitoring tools such as Prometheus, Grafana or equivalent platforms.
- Strong understanding of production support, troubleshooting, performance tuning and failure recovery.
- Ability to clearly explain technical design decisions, performance trade-offs and production engineering challenges.
- Strong collaboration skills with Data Science, ML, Backend Engineering, Architecture and Infrastructure teams.
Benefits
The company offers a range of employee benefits including:
- Cashless medical insurance for employees, spouses, and children
- Accidental insurance coverage
- Life insurance coverage
- Retirement benefits including Provident Fund (PF) and Gratuity
- ESI*
- Complementary meal coupons
- Company-paid transportation
- Sodexo benefits for income tax savings
- Paternity & Maternity Leave Benefit
- National Pension Saving
- EL encashment
- Sick Leave
Similar roles
-
Senior Data Engineer R&D 26-27 H/F
MAZARS Levallois-Perret, Ile-de-France, France
-
Data Engineer Intern (Summer 2027)
Lyft Toronto, Ontario, Canada · CA$83K–CA$94K/yr
-
Senior Data Engineer
Prenuvo United States · $90K–$100K/yr
-
Data Engineer
Ford Motor Company Dearborn, Michigan, United States · $115K–$193K/yr
-
Data Engineer, Ground Network Engineering (Gateway)
SpaceX Redmond, Washington, United States · $125K–$200K/yr
-
Senior Data Engineer
Adyen San Francisco, California, United States · $198K–$293K/yr