Synthlane Technologies Private Limited

Data Engineer – Mid & Senior Level

Synthlane Technologies Private Limited

IT Services and IT Consulting · 11-50 employees

Yesterday
Remote data-engineer Senior (5-10 yrs) Full-time
Log in to apply, save this posting, or score it against your profile with AI.

About the role

You will design, develop, and maintain production-grade data pipelines on AWS to transform operational data into secure, AI-ready datasets. The role involves implementing data privacy transformations, orchestration, and comprehensive monitoring to ensure high-quality data delivery.

What they look for

Python SQL AWS Data Engineering Data Pipelines ETL/ELT Apache Spark Parquet Airflow Dagster AWS Step Functions Data Orchestration Data Quality Schema Management Data Transformation CI/CD

Requirements

Candidates must have strong proficiency in Python and SQL, along with hands-on experience in AWS data services and distributed processing technologies like Apache Spark. A solid background in software engineering best practices, including Git, automated testing, and production troubleshooting, is required.

Full description

This is a remote position.

Role Overview

We are looking for Data Engineers at Senior and Mid-Level to join our team in building a privacy-preserving data platform where data engineering meets production-grade software engineering.

You will work on designing, developing, and maintaining reliable data pipelines that transform operational data into high-quality, secure, and AI-ready datasets.

Key Responsibilities

  • Build and maintain production-grade data pipelines on AWS.
  • Extract, transform, validate, and curate large-scale Parquet datasets.
  • Implement data de-identification, masking, and privacy-preserving transformations.
  • Design and maintain data pipeline orchestration, scheduling, retries, and backfill mechanisms.
  • Implement comprehensive data quality checks, monitoring, and alerting.
  • Work with workflow orchestration tools such as Airflow, Dagster, or AWS Step Functions.
  • Contribute to CI/CD pipelines and Infrastructure as Code (IaC) practices.
  • Manage schema evolution and schema drift across data sources and pipelines.
  • Provide production support, troubleshooting, and root cause analysis for data pipeline issues.
  • Maintain data catalogs, metadata, and data lineage.
  • Follow software engineering best practices including Git, code reviews, automated testing, and maintainable code.
  • Build reliable and idempotent data pipelines capable of handling retries and large-scale backfills.

Required Skills & Experience

  • Strong proficiency in Python and SQL.
  • Hands-on experience with AWS data services and production data pipelines.
  • Experience with Apache Spark or equivalent distributed data processing technologies.
  • Practical experience with Airflow, Dagster, AWS Step Functions, or similar orchestration tools.
  • Strong understanding of data pipeline architecture, ETL/ELT, and data transformation.
  • Experience working with Parquet and large-scale datasets.
  • Understanding of data quality, schema management, monitoring, and alerting.

Strong software engineering practices including:

  • Git and version control
  • Code reviews
  • Automated testing
  • Idempotency
  • Error handling
  • Retries and backfills
  • Experience supporting and troubleshooting production data pipelines.
  • Ability to work effectively with cross-functional engineering and data teams.

Requirements

Required Skills

Python | SQL | AWS | Data Engineering | Data Pipelines | ETL/ELT | Apache Spark | Parquet | Airflow | Dagster | AWS Step Functions | Data Orchestration | Data Quality | Schema Management | Data Transformation | Production Support | Git | CI/CD | Automated Testing | Data De-identification | Data Lineage | Data Catalog

Good to Have

Debezium | AWS DMS | Apache Iceberg | Delta Lake | Apache Hudi | Data Masking | Data Tokenization | Terraform | CloudFormation | ML/AI Training Data | Privacy-Preserving Data

Similar roles