Databricks Data Engineer India
Cogniify Pune, Maharashtra, India
Business Consulting and Services · 11-50 employees
About the role
The role involves supporting, maintaining, and optimizing existing Databricks-based data applications and pipelines to ensure production stability. You will collaborate with engineering teams to troubleshoot issues, manage workspace configurations, and implement data quality improvements.
What they look for
Requirements
Candidates must have 6–9 years of experience in data engineering with strong proficiency in Databricks, Apache Spark, and SQL. A solid understanding of Delta Lake concepts, cloud platforms, and production data pipeline support is required.
Full description
Job Title: Databricks Data Engineer
Location: Bangalore, Hyderabad , Pune
Work Hours: IST
Experience: 6–9 years
Role Overview
We are looking for an experienced Databricks Data Engineer to support, maintain, and enhance existing Databricks-based data applications and pipelines. The role focuses on ensuring reliability, performance, and scalability of production Databricks workloads rather than building net-new platforms from scratch. You will work closely with data, analytics, and engineering teams to keep critical data applications stable, optimized, and aligned with business needs.
Key Responsibilities
Support and maintain existing Databricks applications, notebooks, jobs, and Delta Lake pipelines in production.
Monitor, troubleshoot, and resolve issues related to job failures, performance degradation, data quality, and cluster utilization.
Optimize existing Spark jobs, SQL queries, and Delta tables for cost, performance, and reliability.
Manage and improve Databricks workspace configurations, including clusters, job scheduling, access controls, and Unity Catalog (where applicable).
Implement and maintain data quality checks, logging, alerting, and basic observability for Databricks workloads.
Collaborate with stakeholders to understand requirements for enhancements or bug fixes on existing applications.
Perform incremental improvements, refactoring, and technical debt reduction on current Databricks solutions.
Ensure adherence to best practices around security, governance, and cost management within the Databricks environment.
Document existing pipelines, dependencies, and operational runbooks.
Participate in on-call or support rotations as needed to maintain production stability (within EST working hours).
Required Qualifications
6–9 years of overall experience in data engineering, with strong hands-on experience in Databricks.
Solid proficiency in Apache Spark (PySpark and/or Scala) and SQL.
Proven experience supporting and optimizing production Databricks workloads (jobs, notebooks, Delta Lake, workflows).
Strong understanding of Delta Lake concepts (ACID transactions, time travel, optimization techniques such as Z-ordering, vacuum, optimize).
Experience with Databricks Job clusters, Interactive clusters, and performance tuning (partitioning, caching, shuffle optimization, autoscaling).
Familiarity with data modeling, ETL/ELT patterns, and production data pipeline support.
Experience working with cloud platforms (preferably Azure, AWS, or GCP) in the context of Databricks.
Ability to troubleshoot complex Spark and Databricks issues independently.
Strong communication skills and ability to work effectively in a remote, EST-aligned team.
Preferred Qualifications
Experience with Unity Catalog, Databricks SQL, or Lakehouse architecture.
Knowledge of CI/CD practices for Databricks (e.g., Databricks Asset Bundles, Git integration, Terraform/ARM templates).
Familiarity with orchestration tools (Airflow, Azure Data Factory, or Databricks Workflows).
Exposure to data quality frameworks, monitoring tools, or cost optimization initiatives on Databricks.
Experience supporting analytics or BI teams consuming Databricks data products.
Work Arrangement
Persistent Office
Similar roles
-
Senior Data Engineer, Economy
Roblox San Mateo, California, United States · $243K–$295K/yr
-
Fabric Senior Data Engineer
EXL Pune, Maharashtra, India
-
Cloud Data Engineer Senior
Paradigma Digital - Nuestras ofertas de Empleo Pozuelo de Alarcón, Community of Madrid, Spain
-
Data Engineer Senior
Paradigma Digital - Nuestras ofertas de Empleo Pozuelo de Alarcón, Community of Madrid, Spain
-
Forward Deployed Data Engineer - Houston
Indicium AI Houston, Texas, United States · $180K–$230K/yr
-
Data Engineer
EXL Atlanta, Georgia, United States