Data Engineer
Kumaran Systems · Toronto, Ontario, Canada
IT Services and IT Consulting · 1,001-5,000 employees
About the role
The role involves building and maintaining end-to-end ETL pipelines, including batch and streaming processes using Databricks. You will be responsible for data modeling, performance tuning, and implementing data governance and security practices.
What they look for
Requirements
Candidates must have over 5 years of experience with Databricks, Apache Spark, and cloud platforms like Azure. Strong proficiency in SQL, data transformation, and building complex data architectures such as CDC and SCD is required.
Full description
Required Skills.
- Strong hands-on experience with Data-bricks and Apache Spark (PySpark/Scala).
- Experience in SQL and data transformation techniques.
- Knowledge of ETL tools and data pipeline development.
- Experience working with cloud platforms (Azure/AWS/GCP)
- Strong Azure cloud background
- Understanding of data warehousing concepts.
- Strong problem-solving and analytical skills.
- Hands-on experience with Azure Data bricks or Delta Lake, in building ETL pipelines : batch (autoloader) and Spark structured streaming
- Knowledge of data modelling and performance tuning in Spark.
- Exposure to CI/CD pipelines and DevOps practices.
- Familiarity with data governance and security practices.
- Strong hands on working experience of Unity catalog:
- Hands on exposure to Creating end to end environments : creating catalogs, schemas, tables . materialized views, functions, volumes
- Experience in building SCD 1 and SCD2 (slowly changing dimensions ) on dimension tables
- Experience in building CDC (change data capture pipelines)
- Strong hands on experience with Lakehouse federation , creating foreign catalogs to get data from external sources
- Strong understanding of databricks partitioning , Liquid clustering