Synechron

Data Engineer – PySpark, Hadoop, Hive & Big Data Pipeline Development

Synechron Bengaluru, Karnataka, India

Technology, Information and Internet · 10,001+ employees

Yesterday
data-engineer Senior (5-10 yrs) Full-time India
Create a free account to apply — email only, no card. You can also save this posting or score it against your profile with AI.

About the role

Design, develop, and maintain scalable data pipelines using PySpark and the Cloudera Data Platform. Collaborate with cross-functional teams to build reliable data ingestion and transformation processes while ensuring data quality and operational reliability.

What they look for

PySpark Cloudera Data Platform Apache Spark Hadoop Hive HDFS ETL Data Warehousing Python SQL Data Pipelines Distributed Computing Data Engineering Big Data Architecture Data Quality Cloud Technologies

Requirements

Requires 6-10 years of professional experience in data engineering or big data development with strong expertise in PySpark, Hadoop, and Hive. A bachelor's degree in a relevant field is required, along with demonstrated experience in distributed data-processing environments.

Benefits

Flexible workplace arrangements Mentoring Internal mobility Learning and development programs

Full description

Job Summary

Synechron is seeking a Data Engineer with strong expertise in PySpark and the Cloudera Data Platform (CDP) to design, develop, and maintain scalable data pipelines.

The role is responsible for building reliable data ingestion and transformation processes, working with large-scale datasets, and supporting distributed data-processing environments. The Data Engineer will contribute to business objectives by improving data availability, quality, integrity, processing efficiency, and operational reliability.

The successful candidate will collaborate with cross-functional teams to deliver data-driven solutions that support business and technology requirements.

Software Requirements

Required Software Skills

  • PySpark: Strong hands-on experience developing, optimizing, and maintaining production-grade data pipelines.
  • Cloudera Data Platform (CDP): Strong practical experience working with CDP environments and related data-processing capabilities.
  • Apache Spark: Hands-on experience with Spark processing, job execution, performance tuning, and troubleshooting.
  • Hadoop: Experience working with Hadoop-based distributed data ecosystems.
  • Hive: Experience developing and optimizing Hive queries, tables, and data-processing workflows.
  • HDFS: Experience managing and processing data stored in the Hadoop Distributed File System.
  • ETL tools and processes: Practical experience designing data ingestion, transformation, validation, and loading workflows.
  • Data warehousing technologies: Experience working with data warehouse concepts, structures, and processing patterns.
  • Cloud and distributed data environments: Experience with cloud-based or distributed data-processing platforms.
  • Version experience: Experience with the versions of PySpark, Spark, Hadoop, Hive, HDFS, CDP, and associated tools adopted by the assigned Synechron project; ability to work with current supported releases and understand version-related compatibility considerations.

Preferred Software Skills

  • Experience with cloud-native data platforms and managed distributed-processing services.
  • Familiarity with workflow orchestration, data- pipeline monitoring, scheduling, and alerting tools.
  • Experience with source-control, code-review, build, deployment, and automated testing tools.
  • Familiarity with data-quality, metadata-management, lineage, and observability tools.
  • Experience working with containerized or infrastructure-as-code environments.
  • Exposure to streaming, real-time processing, or event-driven data platforms.

Equivalent tools may be considered where they provide comparable capabilities.

Overall Responsibilities

  • Design and develop scalable data pipelines using PySpark and Cloudera Data Platform (CDP).
  • Build, optimize, and maintain data ingestion and transformation processes for large-scale datasets.
  • Develop reusable, maintainable, and testable data-processing components.
  • Implement data validation and quality checks to support data accuracy, consistency, integrity, and availability.
  • Work with Hadoop, Hive, HDFS, Spark, ETL processes, and related big data technologies.
  • Apply data warehousing and big data architecture principles when designing data solutions.
  • Collaborate with cross-functional teams to understand data requirements and deliver data-driven solutions.
  • Monitor data workflows, processing jobs, resource utilization, and pipeline performance.
  • Investigate and resolve data-quality issues, job failures, performance bottlenecks, and operational incidents.
  • Optimize Spark jobs, PySpark code, queries, data layouts, partitioning, and resource usage where required.
  • Document data pipelines, transformation logic, dependencies, operational procedures, and known limitations.
  • Support deployment, release, testing, and production-readiness activities for data solutions.
  • Contribute to improvements in data engineering standards, development practices, monitoring, and operational support.
  • Consider sustainable engineering practices by reducing unnecessary data movement, optimizing compute and storage usage, reusing pipeline components, and supporting maintainable solutions.
  • Deliver data pipelines that meet agreed requirements for quality, performance, reliability, security, and availability.
  • Escalate material risks, dependencies, data-quality concerns, and delivery constraints through appropriate channels.

Technical Skills (By Category)

Programming Languages

Essential

  • Strong hands-on experience with Python, particularly for PySpark-based data processing.
  • Experience writing maintainable, modular, testable, and performance-conscious data-processing code.
  • Ability to use SQL for data querying, validation, transformation, and analysis.

Preferred

  • Familiarity with Scala or another language used in distributed data processing.
  • Experience writing shell scripts for data operations, job execution, monitoring, or automation.

Databases/Data Management

Essential

  • Hands-on experience with Hive, HDFS, Hadoop, and data warehousing concepts.
  • Understanding of structured, semi-structured, and large-scale distributed data.
  • Experience designing data ingestion, transformation, storage, and retrieval processes.
  • Knowledge of data quality, data integrity, validation, partitioning, schema management, and data availability.
  • Understanding of ETL processes and big data architecture.

Preferred

  • Experience with relational and non-relational databases.
  • Familiarity with data lakes, lakehouse architectures, metadata management, data lineage, or data cataloguing.
  • Experience working with incremental loads, change-data processing, historical data, and schema evolution.
  • Exposure to batch and streaming data-processing patterns.

Cloud Technologies

Essential

  • Experience working in cloud or distributed data-processing environments.
  • Understanding of cloud-related considerations such as scalability, storage, compute usage, availability, access controls, and operational monitoring.
  • Ability to design or support data pipelines that run reliably in distributed environments.

Preferred

  • Experience with cloud-native data platforms, managed Spark services, object storage, or cloud-based data warehouses.
  • Familiarity with cloud monitoring, logging, automation, and cost-optimization practices.
  • Experience migrating or modernizing Hadoop or CDP workloads in cloud environments.

Frameworks and Libraries

Essential

  • Strong experience with PySpark and Apache Spark.
  • Experience using Spark DataFrame, SQL, transformation, action, partitioning, and optimization capabilities.
  • Ability to select appropriate Spark processing patterns for scalability, reliability, and performance.
  • Experience working with ETL frameworks or reusable data-pipeline components.

Preferred

  • Familiarity with structured streaming or other distributed streaming frameworks.
  • Experience with data-validation, testing, serialization, or data-format libraries.
  • Exposure to reusable pipeline frameworks and configuration-driven processing.

Development Tools and Methodologies

Essential

  • Experience developing, testing, deploying, monitoring, and supporting data pipelines across development and production environments.
  • Familiarity with source control, code reviews, defect tracking, and controlled release practices.
  • Experience troubleshooting failed jobs, data-quality issues, processing delays, and resource-related problems.
  • Understanding of batch processing, distributed computing, ETL development, and software development lifecycle practices.
  • Ability to document technical solutions, operational procedures, dependencies, and data-processing logic.

Preferred

  • Experience with workflow orchestration and scheduling tools.
  • Familiarity with continuous integration and continuous delivery practices.
  • Experience with automated testing, data-pipeline observability, alerting, and performance dashboards.
  • Exposure to agile development practices and iterative delivery.

Security Protocols

Essential

  • Understanding of secure data handling, access control, authentication, authorization, and protection of sensitive information.
  • Ability to apply security and privacy requirements when designing, processing, storing, and transferring data.
  • Awareness of secure coding, credential management, auditability, and least-privilege principles.

Preferred

  • Experience implementing security controls within Hadoop, Hive, HDFS, CDP, cloud, or distributed data environments.
  • Familiarity with encryption, data masking, secure data transfer, and role-based access controls.
  • Experience supporting data-governance, compliance, or audit requirements.

Experience Requirements

  • 6–10 years of professional experience in data engineering, big data engineering, ETL development, or a related role.
  • Strong production experience with PySpark and Cloudera Data Platform (CDP).
  • Hands-on experience with Hadoop, Hive, HDFS, Spark, ETL processes, data warehousing, and big data architecture.
  • Experience designing, developing, optimizing, monitoring, and troubleshooting scalable data pipelines.
  • Experience working with large-scale datasets and distributed data-processing environments.
  • Experience collaborating with cross-functional teams to translate data requirements into reliable technical solutions.
  • Preferred experience in technology-focused, data-intensive, financial, analytical, or other environments requiring reliable and governed data processing.
  • Candidates may also qualify through equivalent combinations of data engineering, ETL development, distributed computing, platform engineering, analytics engineering, or data-platform support experience that demonstrate the required capabilities and outcomes.

Day-to-Day Activities

  • Design, develop, test, and optimize PySpark data pipelines and CDP-based ingestion and transformation workflows.
  • Collaborate with engineering, analytics, business, and operational teams to clarify data requirements, dependencies, delivery timelines, and implementation options.
  • Monitor scheduled and production workflows, investigate failures or performance issues, and deliver fixes, enhancements, documentation, and operational updates.
  • Make technical decisions within the agreed solution scope regarding data-processing logic, performance improvements, validation controls, and pipeline reliability; escalate material risks when required.

Qualifications

  • Bachelor’s degree in Computer Science, Engineering, Information Technology, Data Engineering, or a related field; equivalent relevant professional experience may be considered.
  • Professional certification in data engineering, cloud technologies, big data, Spark, Hadoop, or a related discipline is preferred but not mandatory.
  • Completion of relevant training in PySpark, CDP, distributed data processing, data warehousing, cloud technologies, data security, or software engineering practices is preferred.
  • Demonstrated experience delivering and supporting production data pipelines in distributed or cloud-based environments is required.
  • Commitment to continuous professional development through relevant technical training, hands-on learning, knowledge sharing, and awareness of evolving data engineering practices.

Professional Competencies

  • Apply structured problem-solving and analytical reasoning to investigate data-quality issues, job failures, performance bottlenecks, and processing risks.
  • Work effectively with cross-functional teams, contribute to technical discussions, and support shared delivery objectives.
  • Communicate data requirements, technical options, dependencies, risks, and delivery updates clearly to technical and non-technical stakeholders.
  • Adapt to changing data volumes, platform capabilities, business requirements, technology versions, and delivery priorities.
  • Identify opportunities to improve pipeline scalability, reliability, reusability, data quality, automation, and efficient use of compute and storage resources.
  • Manage development tasks, operational priorities, dependencies, documentation, and delivery commitments while maintaining attention to data integrity and solution quality.

S​YNECHRON’S DIVERSITY & INCLUSION STATEMENT  

Diversity & Inclusion are fundamental to our culture, and Synechron is proud to be an equal opportunity workplace and is an affirmative action employer. Our Diversity, Equity, and Inclusion (DEI) initiative ‘Same Difference’ is committed to fostering an inclusive culture – promoting equality, diversity and an environment that is respectful to all. We strongly believe that a diverse workforce helps build stronger, successful businesses as a global company. We encourage applicants from across diverse backgrounds, race, ethnicities, religion, age, marital status, gender, sexual orientations, or disabilities to apply. We empower our global workforce by offering flexible workplace arrangements, mentoring, internal mobility, learning and development programs, and more.

All employment decisions at Synechron are based on business needs, job requirements and individual qualifications, without regard to the applicant’s gender, gender identity, sexual orientation, race, ethnicity, disabled or veteran status, or any other characteristic protected by law.

Candidate Application Notice

Similar roles