Architect - Data Engineer
TransUnion Hyderabad, Telangana, India
IT Services and IT Consulting · 10,001+ employees
About the role
Design and implement scalable, resilient data processing solutions using Apache Spark and cloud-native technologies. Provide technical leadership and mentorship to engineering teams while ensuring alignment with enterprise architecture and security standards.
What they look for
Requirements
Requires 10+ years of experience in software or data engineering with deep expertise in Apache Spark and distributed systems. Proficiency in Scala, Java, or Python and experience with cloud platforms like AWS or GCP is essential.
Full description
TransUnion's Job Applicant Privacy Notice
Team Overview
The team is focused on building and evolving job manager and distributed processing solutions using Apache Spark and cloud technologies.
This is a hybrid position and involves regular performance of job responsibilities virtually as well as in-person at an assigned TU office location for a minimum of two days a week. Role Overview And Core Responsibilities
Architecture & Solution Design
Responsibilities
Design scalable and resilient data processing solutions using Apache Spark and modern data platform technologies. Create architecture designs, reference patterns, and technical specifications for distributed data platforms. Guide engineering teams on architecture, design, and implementation decisions. Ensure solutions align with enterprise architecture standards, security requirements, and engineering best practices. Evaluate technology choices, frameworks, and platform capabilities to address evolving business needs. Identify architectural risks and recommend mitigation strategies. Contribute to platform modernization and cloud transformation initiatives.
Apache Spark & Data Platform Leadership
Responsibilities
Provide technical leadership for Apache Spark-based batch and streaming data processing solutions. Guide development teams on Spark architecture, performance optimization, and operational best practices. Define design patterns for scalable, fault-tolerant, and maintainable Spark applications. Assist teams in resolving complex technical challenges involving distributed processing and large-scale data workloads. Drive performance improvements through effective partitioning, query optimization, resource management, caching, and tuning strategies. Establish standards for observability, monitoring, reliability, and operational excellence.
Solution Architecture & Engineering Guidance
Responsibilities
Collaborate with engineering teams throughout the software development lifecycle. Conduct architecture and design reviews for new features, data pipelines, and platform enhancements.
Support engineering teams in implementing scalable and maintainable solutions.
Promote adoption of CI/CD, automated testing, infrastructure-as-code, and DevOps best practices. Contribute to technical decision-making for platform enhancements and modernization initiatives.
Stakeholder Collaboration
Responsibilities
Work closely with product managers, engineering leads, data scientists, and platform teams. Translate business requirements into scalable technical solutions and architecture designs. Communicate architecture decisions, technical trade-offs, and implementation approaches to stakeholders. Participate in roadmap discussions to ensure alignment between business goals and technology strategy. Collaborate across teams to drive successful implementation of data platform capabilities.
Technical Mentoring
Responsibilities
Mentor engineers on distributed systems design, Spark development, and data engineering best practices. Provide guidance through architecture reviews, design discussions, and technical workshops.
Share knowledge on emerging data platform technologies and architectural patterns.
Required Knowledge And Experiences
Process & Quality
Responsibilities
Ensure architecture and design activities follow established SDLC and governance standards. Participate in code reviews, design reviews, and architecture assessments. Promote reliability, maintainability, scalability, and performance considerations throughout solution development. Drive continuous improvement in architecture practices and engineering standards.
Required Knowledge & Experience
Technical Expertise
10+ years of experience in software engineering, data engineering, or distributed systems development.
Strong hands-on expertise with Apache Spark (Spark SQL, Structured Streaming, DataFrames, Dataset APIs). Experience designing and building large-scale distributed data processing systems. Strong knowledge of Spark optimization techniques including:
- Partitioning Strategies
- Shuffle Optimization
- Join Optimization
- Memory Management
- Resource Utilization
- Performance Tuning
Proficiency in Scala, Java, or Python. Experience with technologies such as:
- Hadoop
- Hive
- Iceberg
- AWS EMR
- AWS Glue
- GCP Dataproc
- BigQuery
Experience designing and operating batch and streaming data pipelines. Understanding of cloud-native architecture and distributed systems principles. Experience deploying Spark solutions on AWS and/or GCP platforms.
Scope & Positioning
Individual contributor architecture role focused on distributed data platforms and Spark-based solutions. Provides architecture leadership and technical guidance across one or more engineering teams. Responsible for solution architecture quality, technology selection, design governance, and technical direction within the domain. Partners with Engineering Managers and Technical Leads to deliver scalable and reliable platform capabilities. Serves as a key technical advisor for Spark architecture, distributed systems design, and cloud-native data platforms.
Preferred Qualifications
- Experience with AWS services such as EMR, Glue, S3, Lambda, EKS, ECS, Step Functions, and CloudWatch.
- Experience with GCP services such as Dataproc, BigQuery, Cloud Storage, Dataflow, Pub/Sub, and Composer.
- Experience managing terabyte-to-petabyte scale data processing environments.
- Experience with Spark Structured Streaming, real-time analytics, and event-driven data processing.
- Contributions to Spark optimization, platform engineering, or open-source data ecosystem projects are a plus.
TransUnion Overview:
At TransUnion, we encourage and are committed to creating a real, positive impact and shared sense of purpose within our Workforce for Good, which empowers our people to grow, innovate and contribute to a better future for our communities and customers. We strive to build an environment where our associates are in the driver’s seat of their professional development— while having access to help along the way. We recognize that success comes when our associates thrive both professionally and personally; that’s why we prioritize work/life flexibility and offer resources for our teams across the globe to collaborate and drive excellence.
Be a part of our Workforce for Good – you’ll work with great people, pioneering products and cutting-edge technology.
TransUnion Job Title
Architect, Applications Programming
Similar roles
-
Principal Data Engineer
Bukuwarung Hyderabad, Telangana, India
-
Data Engineer
Bukuwarung South Jakarta, Java, Indonesia
-
Data Engineer
Pixelogic Media Partners, LLC Cape Town, Western Cape, South Africa
-
Data Engineer
Serko Ltd Bengaluru, Karnataka, India
-
Lead Data Engineer
Eltropy Inc. India
-
Senior Data Engineer- Managed services
Telefonica Tech pune, Maharashtra, India