Manila Recruitment

Lead Data Engineer (Agentic) - #35289

Manila Recruitment Muntinlupa, National Capital District, Philippines

Staffing and Recruiting · 11-50 employees

7 h ago
Remote data-engineer Senior (5-10 yrs) Full-time Philippines
Create a free account to apply — email only, no card. You can also save this posting or score it against your profile with AI.

About the role

You will own the architecture and technical direction of a cloud-native data platform while leading a team of data engineers to build scalable, production-grade pipelines. The role involves setting engineering standards, implementing data governance, and guiding the integration of AI coding agents into data workflows.

What they look for

Data Engineering Azure Microsoft Fabric Data Architecture Medallion Architecture Python PySpark Power BI SQL Data Modelling CI/CD Data Governance Data Quality Observability Leadership Agentic Engineering

Requirements

Candidates must have at least 8 years of hands-on data engineering experience and 2 years in a leadership capacity. Proficiency in Microsoft Fabric, Medallion architecture, and end-to-end data pipeline development is required.

Full description

As a Lead Data Engineer, you will own the data engineering architecture and technical direction of our client's Gen 3.0 cloud-native data platform on Azure, defining scalable and reusable patterns for ingestion, transformation, data modelling, and analytics across products. This is a hands-on build and leadership role where you will develop the first working versions of solutions, build production-scale pipelines across the Medallion structure from Bronze to Silver to Gold, and ensure reliable delivery through to Power BI. You will set engineering standards, oversee data quality and governance, own platform performance, reliability and cost, and guide the effective use of AI coding agents in data work. You will also lead a small team of data engineers, guide solution design, review their work, resolve complex technical issues, and establish practices that enable the team to build and operate a reliable, reusable and scalable data platform.

Duties and Responsibilities:

Data Architecture and Platform Ownership: Own the architecture for ingestion, transformation and analytics delivery. Decide the patterns for pipelines, storage layout and workspace structure, and hold them across products so we end up with one platform rather than several. Decisions are recorded with the reasoning and revisited on evidence

Measures

  • A published architecture that every new pipeline is built against
  • New product data onboarded onto the existing pattern rather than a bespoke build
  • Platform decisions recorded with their reasoning and reviewed as things change

Standards, and How Agents Are Used on Data Work:

Write down the standards the team and the tooling both build from: schema conventions, transformation patterns, naming, and what qualifies as fit for reporting. Prove each standard with a working reference implementation. Set how coding agents are used on data work: what is specified first, what is always checked, and what is never taken on trust

Measures

  • A data engineering playbook in the repository covering pipelines, naming and monitoring
  • A reference implementation for every pattern the standard requires
  • Generated transformation logic verified against known data before release, by a check that runs

Engineering Depth and Data Design: Pipelines and transformations are production code. They are designed, reviewed, tested and maintained, not assembled in a portal and left. Idempotency and safe replay, correct incremental logic, modelling for how the data will be queried rather than how it arrives, and clean separation between layers

Measures

  • Design decisions defended at review on principle rather than preference
  • Pipelines that replay and backfill correctly by design, not by luck
  • Models other engineers extend without having to rebuild them

Running the Platform: Deployment, Monitoring and Cost: Everything reaches production through source control and a pipeline, with versioning, monitoring and a recovery path that has been tested rather than assumed. Own what the platform costs as volume grows, including the compute the agentic way of working consumes

Measures

  • Pipelines deployed from source control rather than by hand, with drift detected
  • Monitoring and alerting on every production pipeline, each with a named owner
  • Cost per workload understood and acted on, and recovery proven by exercise

Data Quality and Governance: Put quality checks and validation at each stage so problems are caught where they enter rather than in a customer facing report. Keep metric definitions consistent across datasets, and work with security and infrastructure on access control, row level security and privacy

Measures

  • Validation at every layer boundary, with failures visible and owned
  • One agreed definition per business metric, used by every dataset that reports it
  • Access control and row level security agreed and evidenced rather than assumed

Analytics Enablement: Architect the semantic layer- models, data marts and shared datasets. Guide analysts and report developers on modelling, performance and reuse, so business users can trust and interpret the numbers consistently

Measures

  • Shared datasets reused across reports rather than duplicated per report
  • Report and refresh performance held to agreed targets as volume grows
  • Analysts able to build without needing an engineer for every change

Leading the Data Function: Provide technical leadership to the data engineers: guide solution design, review work, resolve the hard problems, and coach the team in lineage, quality and observability. You work in a team that spans more than one location and lead it as a peer rather than through a chain. You will be judged substantially on what the engineers around you can do because you were here

Measures

  • Named engineers visibly more capable, with review load moving off you over time
  • A plan for scaling the workload and onboarding further engineers
  • Peer assessment from engineers and architects outside your reporting line
  • 8+ years of hands-on data engineering experience, with strong data engineering fundamentals and experience designing and building data solutions end-to-end
  • At least 2 years of experience in a Lead Data Engineer, Technical Lead, or people leadership/management role, with the ability to provide technical direction while remaining hands-on
  • Strong understanding of data architecture and data modelling, including dimensional modelling and layered designs such as Medallion Architecture (Bronze → Silver → Gold), with proven experience implementing these in a project
  • Hands-on experience with Microsoft Fabric is required. Candidates with strong Databricks, Snowflake, or Synapse backgrounds are also credible, provided they have the ability to work with Fabric
  • Proven experience building production-scale data pipelines and transformations from scratch, including ingestion from operational source systems, advanced SQL, incremental loads, and Python or PySpark
  • End-to-end data engineering experience, from source ingestion and pipeline development through data modelling and the Power BI reporting layer
  • Working knowledge of Power BI, including dataflows, dataset design, semantic modelling, and DAX
  • Experience with production ownership of data platforms, including monitoring, refresh reliability, recovery, performance, and cost
  • Knowledge of data governance, quality, observability, lineage, validation, access control, row-level security, and privacy, as well as CI/CD for data using Azure DevOps, GitHub Actions, or Fabric Git integration
  • Strong technical judgement and problem-solving skills, with the ability to independently determine the appropriate technical approach, explain the reasoning behind it, and take ownership through to resolution

Advantageous Skills:

  • Knowledge and awareness of agentic engineering and the ability to understand where these tools add value and where their output needs to be validated against real data. Hands-on experience is a plus

Similar roles