Veltris

Senior Data Engineer

Veltris

IT Services and IT Consulting · 501-1,000 employees

Aug 11
Remote data-engineer Senior (5-10 yrs) Full-time
Log in to apply, save this posting, or score it against your profile with AI.

About the role

The Senior Data Engineer will own and optimize the full ingestion-to-consumption data pipeline using Microsoft Fabric, PySpark, and Azure Data Factory. They are responsible for implementing event-driven ingestion, maintaining data quality frameworks, and managing CI/CD workflows to support business intelligence and AI model feeds.

What they look for

Microsoft Fabric Azure Data Factory PySpark Delta Lake Python SQL KQL Azure DevOps Power BI Data Engineering CDC ETL Data Quality Git Azure Service Bus Data Modeling

Requirements

Candidates must have 5+ years of experience in data engineering with strong proficiency in Microsoft Fabric, Delta Lake, Python, and SQL. The role requires expertise in building production-grade pipelines, managing schema evolution, and utilizing Azure cloud services for scalable data solutions.

Full description

This is a remote position.

Data Engineering & Platform

  • Medallion pipeline ownership — Build, maintain, and optimize Bronze, Silver, and Gold layer pipelines in Microsoft Fabric using ADF (Azure Data Factory) and PySpark notebooks. Own the full ingestion-to-consumption chain for all source systems including Acumatica ERP, HubSpot CRM, AutoStore, RingCentral, website analytics, and marketplace channel data (Net32, Amazon, eBay, and others).
  • Event-driven ingestion — Implement and maintain event-triggered ADF pipelines using Acumatica SaaS push notifications routed through Azure Service Bus and Azure Function App receivers. Build OData watermark fallback pipelines with LastModifiedDateTime-based incremental load, pagination handling, and session concurrency management.
  • CDC and incremental load — Design and manage Change Data Capture strategies including CDF (Change Data Feed) versioning in the watermark registry, last-processed version tracking, and initial vs. incremental load logic. Ensure Silver batching strategy (currently 37 tables across 3 batches) runs within SLA and scale intelligently as new sources are onboarded.
  • Bronze-to-Gold transformation — Develop and maintain Silver transformation notebooks including deduplication, null handling, type casting, and DQ rule application. Build and optimize Gold KPI layer notebooks following a star schema — fact and dimension tables — with orchestration notebooks managing execution sequence and dependency management.
  • Schema management and evolution — Validate Bronze Delta table schemas against source system payloads. Implement schema evolution handling via the Delta transaction log. Maintain the source_object_config and watermark_registry control tables that govern what is ingested and how.
  • Data quality enforcement — Implement and maintain DQ rules (DQ-001 through DQ-020) covering completeness, uniqueness, validity, consistency, timeliness, and referential integrity across all Silver layer tables for both Acumatica and HubSpot sources. Route failed records to the DQ failed records table and ensure clean records only proceed to Gold.
  • CI/CD and environment management — Manage the Git promotion flow from dev to production across Bronze, Silver, and Gold layers using Azure DevOps. Maintain environment parity between dev-platform and prod environments including library versions, environment variables, and connection configurations. Follow the pipeline-based promotion workflow — never direct workspace-to-Git sync.
  • Platform observability — Instrument notebooks with the custom SDK logging framework (send_log, send_metric, send_error) using Azure Log Analytics as the destination. Ensure all production notebooks call logging_setup at the start of each run. Monitor pipeline health dashboards for error rates, duration anomalies, and row count deviations.
  • Pecan AI and ML model data feeds — Maintain the data pipeline feeds that power Pecan AI forecasting and recommendation models including demand forecasting, Vegas inventory forecasting, customer churn propensity, and next-best-product models. Ensure Fabric SQL endpoint bidirectional integration with Pecan is healthy and model retraining pipelines are reliable.

Requirements

Technical Skills & Experience

  • Experience: 5+ years in data analysis, data engineering, or a combined analytics and engineering role, preferably in a fast-paced ecommerce or B2B distribution environment.
  • Microsoft Fabric: Hands-on experience with Fabric Lakehouses, Warehouses, ADF pipelines, PySpark notebooks, Delta Lake (Bronze/Silver/Gold medallion architecture), and Power BI semantic models. Familiarity with Fabric CI/CD and workspace management is a strong advantage.
  • Data engineering: Demonstrated experience building and maintaining production-grade ingestion pipelines, incremental load strategies (CDC, watermark, OData), schema management, and orchestration. Experience with event-driven architectures (Azure Service Bus, Azure Function Apps, Event Hubs) is a plus.
  • Delta Lake: Working knowledge of Delta table format, transaction log, schema enforcement, schema evolution, time travel, VACUUM, and CDF (Change Data Feed) for incremental load tracking.
  • SQL & Python: Expertise in writing complex SQL queries for data extraction and transformation. Proficiency in Python for data engineering and analysis using pandas, PySpark, NumPy, and scripting ETL workflows.
  • KQL: Experience writing KQL queries for Azure Log Analytics, Microsoft Fabric real-time analytics, or Azure Data Explorer dashboards.
  • Cloud platform: Hands-on experience with Azure (ADF, Key Vault, Service Bus, Function Apps, Log Analytics, DevOps) and/or AWS. Familiarity with Microsoft Fabric F-SKU capacity management and FinOps cost governance is a plus.
  • Data visualization: Hands-on experience with Power BI including semantic model design, DAX measures, and building certified dashboards for executive and operational audiences.
  • Data governance: Familiarity with Microsoft Purview, data cataloging, lineage, sensitivity classification, and KPI certification workflows. Knowledge of data quality frameworks and DQ rule implementation.
  • Version control & CI/CD: Proficiency with Git and Azure DevOps for source control, branching strategy, and pipeline-based promotion across dev and production environments.
  • Strategic mindset: Ability to translate high-level business objectives into analytical questions and engineering requirements. Demonstrated skill in communicating findings and technical trade-offs to both technical and non-technical stakeholders including executive-level audiences.

Similar roles