Senior Data Engineer (Java | Apache Spark | AWS)
Synthlane Technologies Private Limited Hyderabad, Telangana, India
IT Services and IT Consulting · 11-50 employees
About the role
Design and develop scalable data-processing applications and production-grade ETL/ELT pipelines using Java and Apache Spark on AWS. Manage data ingestion, transformation workflows, and perform troubleshooting for production performance issues.
What they look for
Requirements
Requires 8+ years of experience with strong hands-on expertise in Java, Apache Spark, and AWS cloud-native data services. Candidates must possess deep knowledge of distributed data processing, SQL, and data modeling.
Full description
Role Overview
We are looking for an experienced Java Spark AWS Data Engineer with strong hands-on expertise in Java, Apache Spark, AWS, SQL, and large-scale data processing.
The ideal candidate will be responsible for designing, developing, and maintaining scalable data pipelines and distributed data-processing applications on AWS. The role requires strong software engineering fundamentals along with practical experience in Spark performance optimization, cloud-native data services, ETL/ELT pipelines, and production support. Candidates must have 8+ years of overall experience, with strong hands-on expertise in Java, Apache Spark, AWS, and large-scale data engineering solutions.
Key Responsibilities
- Design and develop scalable data-processing applications using Java and Apache Spark.
- Build and maintain production-grade ETL/ELT data pipelines on AWS.
- Develop distributed batch and real-time data-processing solutions.
- Process large volumes of structured, semi-structured, and unstructured data.
- Build Spark applications using Java, Spark SQL, and DataFrame APIs.
- Optimize Spark jobs for performance, memory utilization, partitioning, and scalability.
- Design and manage data ingestion and transformation workflows using AWS services.
- Work with AWS services such as Amazon S3, EMR, Glue, Lambda, Athena, Redshift, RDS, and CloudWatch.
- Develop and integrate REST APIs and backend services using Java where required.
- Implement data validation, reconciliation, quality checks, and error-handling mechanisms.
- Design scalable data models and data warehouse solutions.
- Troubleshoot Spark jobs, data pipeline failures, and production performance issues.
- Implement monitoring, logging, and alerting for data-processing workloads.
- Participate in system design, architecture discussions, and code reviews.
- Write unit, integration, and data-pipeline tests.
- Support CI/CD pipelines and automated deployments.
- Collaborate with Data Engineers, Architects, DevOps/SRE teams, and business stakeholders.
- Perform root-cause analysis and implement permanent fixes for production issues.
Requirements
Required Skills – Comma-Separated
Java, Java 8, Java 11, Java 17, Apache Spark, Spark SQL, Spark DataFrames, Distributed Data Processing, ETL, ELT, Data Engineering, Data Pipelines, Data Ingestion, Data Transformation, Data Validation, Data Quality, SQL, AWS, Amazon S3, Amazon EMR, AWS Glue, Amazon Athena, Amazon Redshift, AWS Lambda, Amazon RDS, AWS CloudWatch, AWS IAM, Batch Processing, Data Warehousing, Data Modeling, Parquet, JSON, CSV, Spark Optimization, Performance Tuning, Partitioning, Broadcast Joins, Caching, Git, Maven, Gradle, CI/CD, Linux, Production Support, Troubleshooting, Root Cause Analysis
Similar roles
-
Lead Java Developer
F24 City of Zagreb, Croatia · €55K–€65K/yr
-
QA Automation Engineer with Java
Veeam Software Lisbon, Portugal
-
Senior Software Engineer (Java) - Banking Accounts
Adyen Chicago, Illinois, United States · $180K–$243K/yr
-
Engineering Manager (Java & Spring Boot)
LegalAndGeneral Brighton and Hove, England, United Kingdom
-
Java Architect
HEXAWARE United States
-
Java Developer 1015
Veracity Ventures Inc. Canada