Sr. Data Engineer - (RealTime Streaming)
GSSTech Group · Bengaluru, Karnataka, India
IT Services and IT Consulting · 201-500 employees
About the role
The candidate will design, develop, and manage high-throughput, low-latency real-time streaming architectures using Kafka and Flink. They are responsible for building scalable data pipelines, optimizing cluster performance, and ensuring system reliability through infrastructure automation.
What they look for
Requirements
The role requires extensive experience in building enterprise-grade real-time data platforms and managing large-scale distributed systems. Proficiency in Kafka, Flink, Java, PySpark, and cloud-native DevOps practices is mandatory.
Full description
We are looking for an experienced Senior Data Engineer with strong expertise in real-time streaming technologies and large-scale data engineering solutions. The ideal candidate will have hands-on experience designing, building, and managing highly scalable and fault-tolerant streaming data platforms using Kafka, Flink, Java, and PySpark.
The candidate will be responsible for developing high-throughput, low-latency data pipelines, managing Kafka clusters, implementing cloud-native data solutions, and ensuring system reliability, security, and scalability in enterprise environments.
Key Responsibilities• Design, develop, implement, and manage Kafka-based real-time streaming architectures capable of handling high-volume and low-latency workloads.
- Build and maintain scalable streaming data pipelines using Kafka ecosystem components, including:
- Kafka Connect
- ksqlDB
- Schema Registry
- Develop and optimize real-time data processing applications using:
- Apache Flink
- Java
- PySpark
- Perform Kafka cluster setup, administration, configuration, tuning, monitoring, and performance optimization.
- Ensure Kafka clusters are highly available and implement disaster recovery strategies and best practices.
- Design and implement fault-tolerant, scalable, and resilient streaming systems for enterprise-grade applications.
- Work with cloud-native applications and modern data engineering frameworks to build scalable solutions.
- Automate infrastructure provisioning and deployment using Infrastructure-as-Code (IaC) tools such as Terraform.
- Implement and follow GitOps practices for deployment automation and infrastructure management.
- Apply robust security standards and best practices, including:
- SSL / mTLS
- SASL authentication
- ACL management
- Collaborate with cross-functional teams to deliver real-time data engineering solutions aligned with business and analytics requirements.
Required Skills & TechnologiesMandatory Skills• Apache Kafka
- Apache Flink
- Java
- PySpark
- Real-Time Streaming Architecture
- Kafka Cluster Management
- Kafka Connect
- ksqlDB
- Schema Registry
- Streaming Data Pipelines
- Cloud-Native Applications
- Terraform
- Infrastructure-as-Code (IaC)
- GitOps
- High Availability & Disaster Recovery
- Performance Tuning & Optimization
- Fault Tolerance & Scalability
- SSL / mTLS
- SASL
- ACLs
Preferred Experience• Strong experience in building enterprise-grade real-time data platforms.
- Hands-on experience working in high-throughput and low-latency environments.
- Experience with large-scale distributed data processing systems.
- Exposure to cloud platforms and modern DevOps practices will be an added advantage.