Senior data engineer
IDBC Kft Budapest, Central Hungary, Hungary
IT Services and IT Consulting · 201-500 employees
Applying here? Try the free cover letter tool — paste this posting and your résumé, no account needed.
About the role
Design and develop scalable ETL/ELT pipelines using Azure Databricks and automate data workflows with Azure Data Factory. Collaborate with cross-functional teams to implement data modeling, security, and machine learning integration while ensuring high-quality data governance.
What they look for
Requirements
Requires 6-8+ years of professional experience in data engineering with mandatory hands-on expertise in Databricks. Candidates must possess strong SQL and Python skills along with experience in Microsoft Azure cloud-based data platforms.
Full description
- Data
Pipeline Design & Development (ETL/ELT)
o Develop ETL/ELT Pipelines: Build scalable ETL/ELT pipelines using Azure Databricks, integrating multiple data sources like Azure Blob Storage, Azure Data Lake, SQL Databases, etc.
o Automation of Data Workflows: Automate data ingestion, transformation, and loading processes through Azure Data Factory (ADF), Databricks workflows, and Azure Functions.
o Real-time Data Processing: Implement real-time streaming data pipelines using Databricks Structured Streaming for use cases such as IoT or event-driven architectures.
- Data
Transformation & Modeling
o Data Cleaning & Transformation: Leverage Apache Spark in Databricks to process large datasets efficiently, performing data cleansing, transformation, and enrichment.
o Data Modeling: Design and implement dimensional data models (e.g., star schema) optimized for performance and querying in Azure Synapse Analytics or other reporting layers.
o Delta Lake Implementation: Use Delta Lake for reliable and scalable ACID-compliant data storage and to optimize data for batch and stream processing.
- Optimization
& Performance Tuning
o Optimize Data Processing: Tune Databricks notebooks and jobs for performance, leveraging Databricks' autoscaling features and optimizing Apache Spark configurations for specific workloads.
o Data Partitioning & Indexing: Implement best practices for partitioning large datasets, managing table storage formats (Parquet, Delta), and indexing data for faster querying.
o Cluster Management: Manage Databricks clusters (autoscaling, sizing, and costs), ensuring efficient resource utilization on Azure.
- Collaboration
& Integration with Azure Services
o Azure Integration: Integrate Databricks with other Azure services like Azure Data Lake Storage, Azure SQL, Azure Synapse, Azure Key Vault (for security), and Power BI for seamless data flow and analysis.
o Monitoring & Alerts: Set up monitoring, alerting, and logging of data pipelines using tools like Azure Monitor, Databricks Jobs, and Azure Log Analytics to ensure smooth operations.
- Data
Governance & Security
o Data Security: Ensure security measures such as data encryption, role-based access control (RBAC), and compliance with GDPR and other regulations using Azure Active Directory (AAD) and Databricks Secrets.
o Data Quality Management: Implement data quality checks in Databricks pipelines, ensuring consistency, accuracy, and validity of data.
o Version Control & CI/CD: Use version control systems like Git and implement CI/CD pipelines using Azure DevOps for Databricks notebooks and data workflows.
- Collaboration
& Documentation
o Cross-functional Collaboration: Collaborate with data scientists, analysts, and business users to develop insights and solutions that drive business objectives.
o Documentation: Maintain thorough documentation of data architectures in DF Confluence, processes, and pipelines in Databricks for future scalability and team collaboration.
- Advanced
Analytics & Machine Learning
o Machine Learning Integration: Collaborate with data science teams to build and deploy machine learning models using Databricks MLflow and integrate with Azure's machine learning services for operationalization.
o Data Exploration: Support exploratory data analysis and business intelligence needs using Databricks Notebooks and integrate with Azure Power BI or other visualization tools.
Requirements
- Hands-on experience with Databricks (mandatory) and proven expertise in designing, developing, and maintaining data solutions on the Databricks platform.
- 6-8+ years of professional experience in Data Engineering
- Strong SQL expertise
- Proficiency in Python for data processing, automation, and backend data development.
- Experience with Microsoft Azure, including cloud-based data services and data platform implementations.
- Ability to work in modern cloud-based data environments and contribute to the migration and adoption of Databricks as a strategic data platform.
- Strong analytical and problem-solving skills with a focus on scalable and high-quality data solutions.
Similar roles
-
Data Engineer
EKN Engineering Irvine, California, United States
-
Senior Workforce Analytics & Data Engineer
RGP San Jose, California, United States · $198K–$229K/yr
-
Data Engineer
Bluetab, an IBM Company San Isidro, Lima, Peru
-
AI Data Engineer
UKG Lowell, Massachusetts, United States · $72K–$103K/yr
-
System Data Engineer
Ericsson Stockholm, Sweden
-
Data Engineer – Cloud & AI Solutions
Sopra Steria Assago, Lombardy, Italy · €33K–€40K/yr