ShyftLabs

Technical Data Engineer (Databricks)

ShyftLabs Coimbatore, Tamil Nadu, India

Business Consulting and Services · 201-500 employees

14 h ago
data-engineer Senior (5-10 yrs) Full-time India
Log in to apply, save this posting, or score it against your profile with AI.

About the role

Design, develop, and maintain scalable ETL/ELT pipelines using Databricks, PySpark, and SQL. Collaborate with cross-functional teams to integrate data from multiple sources and ensure data quality through rigorous monitoring and optimization.

What they look for

Python PySpark SQL Databricks AWS REST API ETL ELT Unity Catalog Delta Lake Medallion Architecture Data modeling Git CI/CD Spark performance tuning Data warehousing

Requirements

Requires 5+ years of data engineering experience with at least 2 years of hands-on Databricks expertise. Candidates must possess strong proficiency in Python, PySpark, SQL, and data modeling concepts.

Benefits

Competitive salary Insurance package Learning and development resources

Full description

Position Overview

We are looking for a Data Engineer with hands-on experience in building scalable data pipelines and data engineering solutions on the Databricks Lakehouse Platform. The ideal candidate should have strong expertise in Python, PySpark, SQL, Databricks, AWS, and REST API integrations for data ingestion, managing large volumes of data, and data export

ShyftLabs is a growing data product company that was founded in early 2020 and works primarily with Fortune 500 companies. We deliver digital solutions built to help accelerate the growth of businesses in various industries, by focusing on creating value through innovation.

\n

Job Responsibilities:Design, develop, and maintain scalable ETL/ELT pipelines using Databricks, PySpark, and SQL. ● Integrate data from multiple sources, including databases, Amazon S3, files, and REST APIs. ● Build data pipelines with Databricks Unity Catalog. ● Implement business logic, data transformations, and dimensional data models. ● Create, schedule, monitor, and optimize Databricks Jobs and Workflows. ● Design and manage Delta Lake tables using Medallion Architecture (Bronze, Silver,Gold). ● Ensure data quality through validations, error handling, logging, and monitoring. ● Optimize Spark workloads for performance, scalability, and reliability. ● Collaborate with cross-functional teams to deliver production-ready data solutions.

Basic Qualification:Strong expertise in Python, PySpark, and Advanced SQL. ● Hands-on experience with the Databricks Lakehouse Platform. ● Good understanding of Unity Catalog, Delta Lake, Databricks Workflows/Jobs, Clusters, Notebooks, Repos, and Medallion Architecture. ● Experience integrating with REST APIs for data ingestion and data export. ● Strong knowledge of ETL/ELT development, batch processing, incremental loading, and data transformation. ● Experience with data modeling (Star Schema, Snowflake Schema, Fact & Dimension tables, SCD concepts). ● Understanding of data warehousing concepts and best practices. ● Experience working with structured and semi-structured data (CSV, JSON, Parquet, Delta). ● Knowledge of partitioning, file optimization, Spark performance tuning, and query optimization. ● Experience with Git and CI/CD best practices

Preferred Qualifications:5+ years of experience in Data Engineering with 2+ years of hands-on Databricks experience. ● Experience with Auto Loader, Spark Declarative pipelines, Kafka, Airflow, or dbt is a plus. ● Databricks certification is an added advantage.

\n

We are proud to offer a competitive salary alongside a strong insurance package. We pride ourselves on the growth of our employees, offering extensive learning and development resources.

Similar roles