Unison Group

Data Engineer (ETL & PySpark)

Unison Group Singapore, Singapore

Business Consulting and Services · 11-50 employees

5 h ago
data-engineer Mid (2-5 yrs) Full-time Singapore
Log in to apply, save this posting, or score it against your profile with AI.

About the role

Design, develop, and maintain scalable ETL/ELT data pipelines using Python and PySpark to process structured and unstructured data. Collaborate with cross-functional teams to deliver data solutions while ensuring data quality, security, and performance optimization.

What they look for

Python PySpark Apache Spark SQL ETL/ELT Data Warehousing Data Modeling Cloud Platforms Linux/Unix Shell Scripting Git CI/CD Data Governance Performance Tuning Problem Solving Communication Skills

Requirements

Requires a Bachelor's or Master's degree in Computer Science or a related field with strong expertise in Python, PySpark, and SQL. Candidates should have experience with data warehousing, cloud platforms, and distributed computing frameworks.

Full description

Job Summary

We are looking for a skilled Data Engineer with strong expertise in Python, PySpark, and SQL to design, develop, and optimize scalable data pipelines and ETL processes. The ideal candidate should have experience working with large-scale datasets, distributed computing frameworks, and cloud-based data platforms while ensuring data quality, reliability, and performance.

Key Responsibilities

  • Design, develop, and maintain scalable ETL/ELT data pipelines using Python and PySpark.
  • Develop high-performance data processing solutions for structured and unstructured data.
  • Write optimized SQL queries, stored procedures, and data transformations.
  • Build and maintain data models, data marts, and data warehouses.
  • Perform data cleansing, validation, and quality checks.
  • Optimize Spark jobs for performance, scalability, and resource utilization.
  • Integrate data from multiple sources including APIs, databases, and cloud storage.
  • Collaborate with Data Analysts, Data Scientists, and Business teams to deliver data solutions.
  • Troubleshoot production issues and perform root cause analysis.
  • Ensure data governance, security, and best engineering practices.
  • Participate in code reviews and maintain technical documentation.

Required Skills

  • Strong experience in Python programming.
  • Hands-on experience with PySpark and Apache Spark.
  • Strong SQL skills with query optimization.
  • Experience with ETL/ELT pipeline development.
  • Good understanding of data warehousing concepts.
  • Experience with relational databases (Oracle, SQL Server, PostgreSQL, MySQL, etc.).
  • Knowledge of Linux/Unix environment and Shell Scripting.
  • Familiarity with Git and CI/CD processes.
  • Strong analytical and problem-solving skills.

Preferred Skills

  • Experience with cloud platforms such as AWS, Azure, or GCP.
  • Knowledge of Databricks, AWS Glue, or EMR.
  • Experience with Airflow or other workflow orchestration tools.
  • Understanding of Delta Lake, Hive, or Hadoop ecosystem.
  • Exposure to Kafka or other streaming technologies.
  • Knowledge of Docker and Kubernetes.
  • Experience with Agile/Scrum methodology.

Qualifications

  • Bachelor's or Master's degree in Computer Science, Information Technology, Engineering, or a related field.

Nice to Have

  • Experience with Snowflake or Databricks.
  • Cloud certifications (AWS, Azure, or GCP).
  • Knowledge of DevOps practices and CI/CD pipelines.

Key Competencies

  • Python Development
  • PySpark
  • Apache Spark
  • SQL & Query Optimization
  • ETL/ELT Development
  • Data Warehousing
  • Performance Tuning
  • Data Modeling
  • Cloud Data Engineering
  • Problem Solving
  • Team Collaboration
  • Communication Skills

Similar roles