S

Data Engineer - Databricks

Sedin Technologies Private Limited Chennai, Tamil Nadu, India

Jul 22
data-engineer Senior (5-10 yrs) Full-time India
Log in to apply, save this posting, or score it against your profile with AI.

About the role

Build and maintain scalable ETL/ELT pipelines using Databricks and AWS infrastructure. Collaborate with data scientists and analysts to manage data ingestion and deliver high-quality analytical datasets.

What they look for

Databricks PySpark Spark SQL Delta Lake AWS Python SQL Terraform CloudFormation Git Jenkins GitHub Actions ETL Data Engineering Data Modeling Performance Tuning

Requirements

Requires over 5 years of experience in data engineering with strong proficiency in PySpark, SQL, and AWS services. Candidates must have expertise in designing Medallion Architecture and implementing CI/CD pipelines.

Full description

Job Description – Data Engineer (AWS + Databricks)

Experience: 5+ Years

Location: Flexible / Hybrid

Key Responsibilities

  • Build scalable ETL/ELT pipelines using Databricks (PySpark, Spark SQL, Delta Lake).
  • Design and implement Medallion Architecture (Bronze, Silver, Gold) for Lakehouse.
  • Develop batch and real-time pipelines using Spark Streaming, AWS Kinesis, or Glue Streaming.
  • Manage data ingestion from diverse sources – RDBMS, APIs, S3, Kafka.
  • Perform data profiling, validation, and quality checks.
  • Optimize Spark jobs, partitions, and cluster configurations for performance.
  • Implement and manage data infrastructure using AWS S3, Glue, Athena, Lambda, Step Functions, Redshift, EMR.
  • Set up CI/CD pipelines (GitHub Actions, CodePipeline, Jenkins).
  • Collaborate with analysts and data scientists to deliver analytical datasets.

Technical Skills

  • Databricks: PySpark, Spark SQL, Delta Lake, Unity Catalog, Job Workflows.
  • AWS: S3, Glue, Athena, Kinesis, Lambda, Step Functions, EMR, Redshift, CloudWatch.
  • Programming: Python (pandas), SQL (window functions, optimization).
  • IaC & DevOps: Terraform / CloudFormation, Git, Jenkins, GitHub Actions.
  • Performance Tuning: Spark optimizations, Z-ordering, partitioning, caching.

Similar roles