Data Engineer (ETL & PySpark)
Unison Group Singapore, Singapore
Business Consulting and Services · 11-50 employees
About the role
Design, develop, and maintain scalable ETL/ELT data pipelines using Python and PySpark to process structured and unstructured data. Collaborate with cross-functional teams to deliver data solutions while ensuring data quality, security, and performance optimization.
What they look for
Requirements
Requires a Bachelor's or Master's degree in Computer Science or a related field with strong expertise in Python, PySpark, and SQL. Candidates should have experience with data warehousing, cloud platforms, and distributed computing frameworks.
Full description
Job Summary
We are looking for a skilled Data Engineer with strong expertise in Python, PySpark, and SQL to design, develop, and optimize scalable data pipelines and ETL processes. The ideal candidate should have experience working with large-scale datasets, distributed computing frameworks, and cloud-based data platforms while ensuring data quality, reliability, and performance.
Key Responsibilities
- Design, develop, and maintain scalable ETL/ELT data pipelines using Python and PySpark.
- Develop high-performance data processing solutions for structured and unstructured data.
- Write optimized SQL queries, stored procedures, and data transformations.
- Build and maintain data models, data marts, and data warehouses.
- Perform data cleansing, validation, and quality checks.
- Optimize Spark jobs for performance, scalability, and resource utilization.
- Integrate data from multiple sources including APIs, databases, and cloud storage.
- Collaborate with Data Analysts, Data Scientists, and Business teams to deliver data solutions.
- Troubleshoot production issues and perform root cause analysis.
- Ensure data governance, security, and best engineering practices.
- Participate in code reviews and maintain technical documentation.
Required Skills
- Strong experience in Python programming.
- Hands-on experience with PySpark and Apache Spark.
- Strong SQL skills with query optimization.
- Experience with ETL/ELT pipeline development.
- Good understanding of data warehousing concepts.
- Experience with relational databases (Oracle, SQL Server, PostgreSQL, MySQL, etc.).
- Knowledge of Linux/Unix environment and Shell Scripting.
- Familiarity with Git and CI/CD processes.
- Strong analytical and problem-solving skills.
Preferred Skills
- Experience with cloud platforms such as AWS, Azure, or GCP.
- Knowledge of Databricks, AWS Glue, or EMR.
- Experience with Airflow or other workflow orchestration tools.
- Understanding of Delta Lake, Hive, or Hadoop ecosystem.
- Exposure to Kafka or other streaming technologies.
- Knowledge of Docker and Kubernetes.
- Experience with Agile/Scrum methodology.
Qualifications
- Bachelor's or Master's degree in Computer Science, Information Technology, Engineering, or a related field.
Nice to Have
- Experience with Snowflake or Databricks.
- Cloud certifications (AWS, Azure, or GCP).
- Knowledge of DevOps practices and CI/CD pipelines.
Key Competencies
- Python Development
- PySpark
- Apache Spark
- SQL & Query Optimization
- ETL/ELT Development
- Data Warehousing
- Performance Tuning
- Data Modeling
- Cloud Data Engineering
- Problem Solving
- Team Collaboration
- Communication Skills
Similar roles
-
Data Engineer
Fastmarkets Sofia, Sofia-City, Bulgaria
-
Data Engineer - Data & Technology, SCD
Inter IKEA Group Älmhult, Skåne County, Sweden
-
Senior Data Engineer
Morgan Advanced Materials Poland
-
Senior Backend Data Engineer
Eleos Health Tel-Aviv, Tel-Aviv District, Israel
-
Senior Data Engineer
Version 1 Mumbai City, Maharashtra, India
-
Senior Data Engineer
Blend360 Hyderabad, Telangana, India