Synthlane Technologies Private Limited

Senior Data Engineer (Java | Apache Spark | AWS)

Synthlane Technologies Private Limited Hyderabad, Telangana, India

IT Services and IT Consulting · 11-50 employees

19 h ago
java Principal (10+ yrs) Full-time India
Log in to apply, save this posting, or score it against your profile with AI.

About the role

Design and develop scalable data-processing applications and production-grade ETL/ELT pipelines using Java and Apache Spark on AWS. Manage data ingestion, transformation workflows, and perform troubleshooting for production performance issues.

What they look for

Java Apache Spark AWS SQL Data Engineering ETL ELT Data Pipelines Distributed Data Processing Spark SQL Data Modeling Amazon EMR AWS Glue Amazon Redshift Performance Tuning CI/CD

Requirements

Requires 8+ years of experience with strong hands-on expertise in Java, Apache Spark, and AWS cloud-native data services. Candidates must possess deep knowledge of distributed data processing, SQL, and data modeling.

Full description

Role Overview

We are looking for an experienced Java Spark AWS Data Engineer with strong hands-on expertise in Java, Apache Spark, AWS, SQL, and large-scale data processing.

The ideal candidate will be responsible for designing, developing, and maintaining scalable data pipelines and distributed data-processing applications on AWS. The role requires strong software engineering fundamentals along with practical experience in Spark performance optimization, cloud-native data services, ETL/ELT pipelines, and production support. Candidates must have 8+ years of overall experience, with strong hands-on expertise in Java, Apache Spark, AWS, and large-scale data engineering solutions.

Key Responsibilities

  • Design and develop scalable data-processing applications using Java and Apache Spark.
  • Build and maintain production-grade ETL/ELT data pipelines on AWS.
  • Develop distributed batch and real-time data-processing solutions.
  • Process large volumes of structured, semi-structured, and unstructured data.
  • Build Spark applications using Java, Spark SQL, and DataFrame APIs.
  • Optimize Spark jobs for performance, memory utilization, partitioning, and scalability.
  • Design and manage data ingestion and transformation workflows using AWS services.
  • Work with AWS services such as Amazon S3, EMR, Glue, Lambda, Athena, Redshift, RDS, and CloudWatch.
  • Develop and integrate REST APIs and backend services using Java where required.
  • Implement data validation, reconciliation, quality checks, and error-handling mechanisms.
  • Design scalable data models and data warehouse solutions.
  • Troubleshoot Spark jobs, data pipeline failures, and production performance issues.
  • Implement monitoring, logging, and alerting for data-processing workloads.
  • Participate in system design, architecture discussions, and code reviews.
  • Write unit, integration, and data-pipeline tests.
  • Support CI/CD pipelines and automated deployments.
  • Collaborate with Data Engineers, Architects, DevOps/SRE teams, and business stakeholders.
  • Perform root-cause analysis and implement permanent fixes for production issues.

Requirements

Required Skills – Comma-Separated

Java, Java 8, Java 11, Java 17, Apache Spark, Spark SQL, Spark DataFrames, Distributed Data Processing, ETL, ELT, Data Engineering, Data Pipelines, Data Ingestion, Data Transformation, Data Validation, Data Quality, SQL, AWS, Amazon S3, Amazon EMR, AWS Glue, Amazon Athena, Amazon Redshift, AWS Lambda, Amazon RDS, AWS CloudWatch, AWS IAM, Batch Processing, Data Warehousing, Data Modeling, Parquet, JSON, CSV, Spark Optimization, Performance Tuning, Partitioning, Broadcast Joins, Caching, Git, Maven, Gradle, CI/CD, Linux, Production Support, Troubleshooting, Root Cause Analysis

Similar roles