IDBC Kft

Senior data engineer

IDBC Kft Budapest, Central Hungary, Hungary

IT Services and IT Consulting · 201-500 employees

Jun 19
data-engineer Senior (5-10 yrs) Full-time Hungary
Create a free account to apply — email only, no card. You can also save this posting or score it against your profile with AI.

About the role

Design and develop scalable ETL/ELT pipelines using Azure Databricks and automate data workflows with Azure Data Factory. Collaborate with cross-functional teams to implement data modeling, security, and machine learning integration while ensuring high-quality data governance.

What they look for

Databricks Azure Data Factory Apache Spark Python SQL ETL/ELT Delta Lake Azure Synapse Analytics Azure Blob Storage Azure Data Lake Data Modeling CI/CD Azure DevOps Machine Learning Data Governance Power BI

Requirements

Requires 6-8+ years of professional experience in data engineering with mandatory hands-on expertise in Databricks. Candidates must possess strong SQL and Python skills along with experience in Microsoft Azure cloud-based data platforms.

Full description

  • Data

Pipeline Design & Development (ETL/ELT)

o Develop ETL/ELT Pipelines: Build scalable ETL/ELT pipelines using Azure Databricks, integrating multiple data sources like Azure Blob Storage, Azure Data Lake, SQL Databases, etc.

o Automation of Data Workflows: Automate data ingestion, transformation, and loading processes through Azure Data Factory (ADF), Databricks workflows, and Azure Functions.

o Real-time Data Processing: Implement real-time streaming data pipelines using Databricks Structured Streaming for use cases such as IoT or event-driven architectures.

  • Data

Transformation & Modeling

o Data Cleaning & Transformation: Leverage Apache Spark in Databricks to process large datasets efficiently, performing data cleansing, transformation, and enrichment.

o Data Modeling: Design and implement dimensional data models (e.g., star schema) optimized for performance and querying in Azure Synapse Analytics or other reporting layers.

o Delta Lake Implementation: Use Delta Lake for reliable and scalable ACID-compliant data storage and to optimize data for batch and stream processing.

  • Optimization

& Performance Tuning

o Optimize Data Processing: Tune Databricks notebooks and jobs for performance, leveraging Databricks' autoscaling features and optimizing Apache Spark configurations for specific workloads.

o Data Partitioning & Indexing: Implement best practices for partitioning large datasets, managing table storage formats (Parquet, Delta), and indexing data for faster querying.

o Cluster Management: Manage Databricks clusters (autoscaling, sizing, and costs), ensuring efficient resource utilization on Azure.

  • Collaboration

& Integration with Azure Services

o Azure Integration: Integrate Databricks with other Azure services like Azure Data Lake Storage, Azure SQL, Azure Synapse, Azure Key Vault (for security), and Power BI for seamless data flow and analysis.

o Monitoring & Alerts: Set up monitoring, alerting, and logging of data pipelines using tools like Azure Monitor, Databricks Jobs, and Azure Log Analytics to ensure smooth operations.

  • Data

Governance & Security

o Data Security: Ensure security measures such as data encryption, role-based access control (RBAC), and compliance with GDPR and other regulations using Azure Active Directory (AAD) and Databricks Secrets.

o Data Quality Management: Implement data quality checks in Databricks pipelines, ensuring consistency, accuracy, and validity of data.

o Version Control & CI/CD: Use version control systems like Git and implement CI/CD pipelines using Azure DevOps for Databricks notebooks and data workflows.

  • Collaboration

& Documentation

o Cross-functional Collaboration: Collaborate with data scientists, analysts, and business users to develop insights and solutions that drive business objectives.

o Documentation: Maintain thorough documentation of data architectures in DF Confluence, processes, and pipelines in Databricks for future scalability and team collaboration.

  • Advanced

Analytics & Machine Learning

o Machine Learning Integration: Collaborate with data science teams to build and deploy machine learning models using Databricks MLflow and integrate with Azure's machine learning services for operationalization.

o Data Exploration: Support exploratory data analysis and business intelligence needs using Databricks Notebooks and integrate with Azure Power BI or other visualization tools.

Requirements

  • Hands-on experience with Databricks (mandatory) and proven expertise in designing, developing, and maintaining data solutions on the Databricks platform.
  • 6-8+ years of professional experience in Data Engineering
  • Strong SQL expertise
  • Proficiency in Python for data processing, automation, and backend data development.
  • Experience with Microsoft Azure, including cloud-based data services and data platform implementations.
  • Ability to work in modern cloud-based data environments and contribute to the migration and adoption of Databricks as a strategic data platform.
  • Strong analytical and problem-solving skills with a focus on scalable and high-quality data solutions.

Similar roles