Senior Data Engineer
Clarivate · Noida, Uttar Pradesh, India
Information Services · 10,001+ employees
About the role
You will design, develop, and maintain scalable data warehouse and Lakehouse solutions while building robust ETL/ELT pipelines for data ingestion and transformation. Additionally, you will partner with architects to define data standards and implement data quality, governance, and observability practices.
What they look for
Requirements
Candidates must have a bachelor's degree and at least 5 years of experience in software or data engineering. Strong proficiency in SQL, Python, and experience with Databricks and AWS cloud data services are required.
Full description
We are looking for a Senior Data Engineer to join our Architecture team in Noida, India. This is an amazing opportunity to help shape modern enterprise data platforms in the Intellectual Property space, with a focus on scalable data warehouse, Lakehouse, data integration, reporting, analytics, and future AI-enabled use cases. We have a great skill set in enterprise architecture, data platform engineering, Databricks-led data warehouse/Lakehouse modernization, and AWS cloud data services, and we would love to speak with you if you have experience with SQL, Python, ETL/ELT pipelines, data modelling, Databricks, and AWS.
About You – experience, education, skills, and accomplishments
- Bachelor’s degree in computer science, engineering, information systems, or equivalent experience.
- Min 5 years of experience in software engineering, data engineering, analytics engineering, or enterprise data platform delivery.
- Hands-on experience building data warehouses, Lakehouse platforms, ETL/ELT pipelines, or large-scale data integration solutions.
- Strong SQL skills, including query optimization, data profiling, and analytical data structures.
- Strong Python or equivalent data processing language experience.
- Good understanding of data modelling, including conceptual, logical, physical, normalized, and dimensional modelling techniques.
- Experience with Databricks-based data warehouse/Lakehouse platforms and AWS cloud data services is preferred.
- Experience implementing data quality, validation, observability, and operational monitoring practices.
- Ability to work with architects and senior stakeholders to convert platform needs into actionable engineering designs.
It would be great if you also have . . .
- Experience with Databricks notebooks, jobs/workflows, Delta Lake, Unity Catalog, performance tuning, and production-grade data pipeline delivery.
- Experience with AWS data services such as Amazon S3, AWS Glue, Amazon Redshift, Amazon Athena, AWS Lake Formation, Amazon EMR, or AWS Lambda.
- Experience with metadata management, data lineage, data cataloguing, or governance tools.
- Experience with APIs, event-driven integration, streaming data, or message-based data exchange.
- Experience in Intellectual Property, Legal Tech, finance, enterprise workflow, or product-platform environments.
- Experience supporting modernization programs, platform consolidation, or enterprise architecture initiatives.
What will you be doing in this role?
You will be responsible for,
- Designing, developing, and maintaining scalable data warehouse, Lakehouse, and data integration solutions.
- Building robust ETL/ELT pipelines for ingestion, transformation, enrichment, and publishing of trusted data assets.
- Supporting batch and near-real-time data processing patterns where appropriate.
- Partnering with architects to define data architecture patterns, standards, implementation approaches, and reusable engineering practices.
- Developing conceptual, logical, and physical data models for enterprise data domains.
- Contributing to canonical data models, domain data models, data contracts, and data interoperability standards.
- Creating curated and reusable data assets that enable reporting, analytics, decision support, and AI/ML use cases.
- Implementing data quality checks, validation frameworks, reconciliation controls, observability, and monitoring practices.
- Applying security, privacy, access control, and governance principles to data platform delivery.
- Following CI/CD, version control, testing, documentation, and release management practices for data solutions.
- Participating in design reviews, contributing to architecture decision records, and mentoring engineers on reusable data engineering patterns.
- Identifying technical debt, data duplication, and opportunities for platform simplification.
About the Team
The Architecture team supports Intellectual Property technology by strengthening architecture-led engineering capability, improving consistency across the data ecosystem, and establishing reusable, governed data assets. The team works closely with Enterprise Architects, Product Teams, Engineering Teams, Data Scientists, and business stakeholders to deliver trusted data foundations for reporting, analytics, operational insight, modernization initiatives, and future AI-enabled product capabilities. The preferred data warehouse/Lakehouse platform is Databricks, and AWS is the preferred cloud platform.
Hours of Work
This is a full-time opportunity with Clarivate, 9 hours per day including 1 hour lunch break.
At Clarivate, we are committed to providing equal employment opportunities for all qualified persons with respect to hiring, compensation, promotion, training, and other terms, conditions, and privileges of employment. We comply with applicable laws and regulations governing non-discrimination in all locations.