WPP

Senior Data Engineer

WPP India

Advertising Services · 10,001+ employees

7 h ago
data-engineer Principal (10+ yrs) Full-time India
Create a free account to apply — email only, no card. You can also save this posting or score it against your profile with AI.

About the role

Design, build, and optimize scalable data lakehouse platforms using Google BigQuery or Databricks. Develop robust ETL/ELT pipelines and manage the data lifecycle from ingestion to governance and performance optimization.

What they look for

Google BigQuery Databricks Python PySpark SQL Apache Spark Data Engineering ETL/ELT Cloud Composer Airflow Data Governance Delta Lake CI/CD Git Data Modeling Multi-cloud

Requirements

Requires 8+ years of experience in data engineering with deep expertise in modern lakehouse platforms and multi-cloud environments (Azure, AWS, GCP). A bachelor's degree in a technical field is preferred, along with strong proficiency in Python, PySpark, and advanced SQL.

Benefits

Hybrid work environment Continuous learning opportunities Career development programs

Full description

WPP is the trusted growth partner for the world’s leading brands.

We unite cutting-edge media intelligence and data solutions, world-class creativity, next-generation production, transformative enterprise solutions and expert strategic counsel in a single company – powered by exceptional talent and our agentic marketing platform, WPP Open, to help our clients navigate change, capture opportunity and deliver transformational growth.

We work with the world's most valuable brands and have global reach across 100+ markets, with deep local expertise.

Our people are the key to our success. We're committed to fostering a culture of creativity, belonging and continuous learning, attracting and developing the brightest talent, and providing exciting career opportunities that help our people grow.

For more information, visit WPP.com.

Why we're hiring:

We are seeking a highly skilled and experienced Senior Data Engineer to join our growing data team. In this critical role, you will be instrumental in designing, building, and optimizing our scalable data lakehouse platform using Google BigQuery or Databricks. You will be a key player in developing robust data pipelines that ingest data from various sources, including Google Analytics 4 (GA4), and transform it into reliable, analysis-ready datasets within the lakehouse environment. This role requires deep expertise in modern lakehouse platforms – Google BigQuery and/or Databricks – together with strong skills in SQL, Python, and Apache Spark (PySpark), along with strong hands-on experience across Azure, AWS, and GCP cloud environments, as our data ecosystem spans multiple cloud platforms. You will be responsible for the entire data lifecycle within the lakehouse, from ingestion and transformation to governance and optimization, ensuring data quality and performance. You should be adept at analyzing performance bottlenecks in Spark jobs and BigQuery workloads, providing enhancement recommendations, and collaborating effectively with both technical and non-technical stakeholders.

What you'll be doing:

  • Design, build, and deploy robust ETL/ELT pipelines within the lakehouse platform (Google BigQuery or Databricks) using SQL, Python, PySpark, and Spark SQL.
  • Implement and manage the Medallion Architecture (Bronze, Silver, Gold layers) using Delta Lake or BigQuery datasets to ensure data quality and progressive data refinement.
  • Leverage native ingestion tooling – such as BigQuery Data Transfer Service, Pub/Sub streaming, or Databricks Auto Loader – for efficient, scalable, and incremental ingestion of data from sources like GA4 into the Bronze layer.
  • Develop, schedule, and monitor complex, multi-task data workflows using Cloud Composer (Airflow), BigQuery scheduled queries, or Databricks Workflows.
  • Optimize BigQuery tables (partitioning, clustering, materialised views) and Spark jobs / Delta Lake tables (using techniques like OPTIMIZE, Z-ORDER, and partitioning) for high performance and cost efficiency.
  • Implement data governance, security, and discovery using Dataplex / BigQuery policy tags or Unity Catalog, including managing access controls and data lineage.
  • Write complex, customized SQL queries to manipulate data and support ad-hoc analytical requests from business teams.
  • Develop strategies for data ingestion from multiple sources, using various techniques including streaming, API consumption, and replication.
  • Document data engineering processes, data models, and technical specifications for the lakehouse platform.
  • Conform to agile development practices, including version control (Git), continuous integration/delivery (CI/CD), and test-driven development.
  • Provide production support for data pipelines, actively monitoring and resolving issues to ensure the continuous flow of critical data.
  • Collaborate with analytics and business teams to understand data requirements and deliver well-modelled, performant datasets in the gold layer

What you'll need:

  • Education: Minimum of a bachelor’s degree in computer science, Engineering, Mathematics, or a related technical field preferred.
  • Experience: 8+ years of relevant experience in data engineering, with a significant focus on building data pipelines on distributed systems.

Engineer's Core Skills:

  • Lakehouse Platform Expertise (Google BigQuery and/or Databricks)
  • BigQuery: Deep, hands-on experience with BigQuery architecture, including partitioning, clustering, materialised views, slot/cost optimisation, and diagnosing query performance using query plans and INFORMATION_SCHEMA.
  • Apache Spark / Delta Lake: Strong experience with Spark architecture, writing and optimising PySpark and Spark SQL jobs, and building reliable pipelines on Delta Lake. Proficient with ACID transactions, time travel, schema evolution, and DML operations (MERGE, UPDATE, DELETE).
  • Data Ingestion: Experience with modern ingestion tools, such as BigQuery Data Transfer Service, Pub/Sub / Dataflow streaming, Databricks Auto Loader, and COPY INTO for scalable file processing.
  • Data Governance: Strong understanding of data governance concepts and practical experience implementing security, lineage, and discovery using Dataplex, BigQuery IAM and policy tags, or Unity Catalog.

Core Engineering & Cloud Skills:

  • Programming: 5+ years of strong, hands-on experience in Python, with an emphasis on PySpark for large-scale data transformation.
  • SQL: 6+ years of advanced SQL experience, including complex joins, window functions, and CTEs.
  • Cloud Platforms: 8+ years of strong, hands-on experience working across all three major cloud platforms – Azure, AWS, and GCP – including expertise in cloud storage (ADLS Gen2, S3, Google Cloud Storage), security and identity management (Azure AD/Entra ID, AWS IAM, GCP IAM), and cloud networking. Proven ability to design, deploy, and support data solutions in multi-cloud environments.
  • Data Modeling: Experience designing star schemas and applying data warehouse methodologies to build analytical models (Gold layer).
  • CI/CD & DevOps: Hands-on experience with version control (Git) and CI/CD pipelines (e.g., GitHub Actions, Azure DevOps) for automating the deployment of BigQuery and Databricks assets.

Tools & Technologies:

  • Primary Data Platform: Google BigQuery or Databricks
  • Cloud Platforms: Azure, AWS, and GCP (strong hands-on experience across all three required)
  • Data Warehouses (Integration): Snowflake
  • Orchestration/Transformation: Cloud Composer (Airflow), Databricks Workflows, dbt (data build tool)
  • Version Control: Git/GitHub or similar repositories
  • Infrastructure as Code (Bonus): Terraform
  • BI Tools (Bonus): Looker or Power BI

Who you are:

You're open: we are inclusive and collaborative; we encourage the free exchange of ideas; we respect and celebrate diverse views. We are open-minded: to new ideas, new partnerships, new ways of working.

You're optimistic: we approach all that we do with confidence: to try the new and to seek the unexpected.

You're extraordinary: We are stronger together: through collaboration we achieve the amazing. We are creative leaders and pioneers of our industry; we provide extraordinary every day.

What we'll give you:

Passionate, inspired people – we champion a culture of people that do extraordinary work

Scale and opportunity – we offer the opportunity to create, influence and deliver projects at a scale that is unparalleled in the industry.

Challenging and stimulating work – unique work and the opportunity to join a group of creative problem solvers.

#LI-Hybrid

We believe the best work happens when we're together, fostering creativity, collaboration, and connection. That's why we’ve adopted a hybrid approach, with teams in the office around four days a week. If you require accommodations or flexibility, please discuss this with the hiring team during the interview process.

WPP is an equal opportunity employer and considers applicants for all positions without discrimination or regard to particular characteristics. We are committed to fostering a culture of respect in which everyone feels they belong and has the same opportunities to progress in their careers.

Please read our Privacy Notice (https://www.wpp.com/en/careers/wpp-privacy-policy-for-recruitment) for more information on how we process the information you provide.

Similar roles