Kumaran Systems

Data Engineer

Kumaran Systems · Toronto, Ontario, Canada

IT Services and IT Consulting · 1,001-5,000 employees

Jul 15
Senior (5-10 yrs) Full-time Canada
Log in to apply, save this posting, or score it against your profile with AI.

About the role

The role involves building and maintaining end-to-end ETL pipelines, including batch and streaming processes using Databricks. You will be responsible for data modeling, performance tuning, and implementing data governance and security practices.

What they look for

Databricks Apache Spark PySpark Scala SQL ETL Azure Data Warehousing Delta Lake Unity Catalog SCD CDC Lakehouse Federation CI/CD DevOps Data Governance

Requirements

Candidates must have over 5 years of experience with Databricks, Apache Spark, and cloud platforms like Azure. Strong proficiency in SQL, data transformation, and building complex data architectures such as CDC and SCD is required.

Full description

Required Skills.

  • Strong hands-on experience with Data-bricks and Apache Spark (PySpark/Scala).
  • Experience in SQL and data transformation techniques.
  • Knowledge of ETL tools and data pipeline development.
  • Experience working with cloud platforms (Azure/AWS/GCP)
  • Strong Azure cloud background
  • Understanding of data warehousing concepts.
  • Strong problem-solving and analytical skills.
  • Hands-on experience with Azure Data bricks or Delta Lake, in building ETL pipelines : batch (autoloader) and Spark structured streaming
  • Knowledge of data modelling and performance tuning in Spark.
  • Exposure to CI/CD pipelines and DevOps practices.
  • Familiarity with data governance and security practices.
  • Strong hands on working experience of Unity catalog:
  • Hands on exposure to Creating end to end environments : creating catalogs, schemas, tables . materialized views, functions, volumes
  • Experience in building SCD 1 and SCD2 (slowly changing dimensions ) on dimension tables
  • Experience in building CDC (change data capture pipelines)
  • Strong hands on experience with Lakehouse federation , creating foreign catalogs to get data from external sources
  • Strong understanding of databricks partitioning , Liquid clustering