Gradera Inc.

Data Engineer

Gradera Inc. · Hyderabad, Telangana, India

Software Development · 11-50 employees

Yesterday
Senior (5-10 yrs) Full-time India
Log in to apply, save this posting, or score it against your profile with AI.

About the role

Design, build, and maintain scalable data pipelines using Databricks, PySpark, and Kafka to power digital twin platforms and AI/ML workloads. Collaborate with cross-functional teams to ensure data quality, governance, and performance optimization across operational systems.

What they look for

Databricks PySpark Delta Lake Kafka Data Engineering Data Pipelines Unity Catalog Delta Live Tables Teradata Data Warehousing SQL Machine Learning Data Governance Operational Modeling Time-series Data Agile

Requirements

Requires 7+ years of hands-on data engineering experience with a strong track record of building production-grade pipelines. Proficiency in Databricks, Delta Lake, and Kafka is essential, along with experience in agile environments and operational data modeling.

Full description

About Gradera

Gradera defines a new category of enterprise transformation called Software-Orchestrated Services™ - where software orchestrates human expertise, digital workers, and enterprise systems to deliver governed outcomes at scale. As an AI Native Services firm, we help enterprises redesign how work gets done across operations, product, engineering, customer experience, data, and enterprise workflows to move beyond fragmented AI pilots and disconnected automation toward measurable business outcomes

 

 

Overview

We are seeking skilled Data Engineers to join our Data & Digital Twin Foundation team. You will design, build, and maintain data pipelines that power digital twin platforms, real-time operational systems, and AI/ML workloads. Working closely with data architects, simulation engineers, and ML teams, you will transform raw operational data into high-quality, governed datasets that drive intelligent decision-making.

 

Key Responsibilities

  • Design, develop, and maintain scalable data pipelines using Databricks, PySpark, and Delta Lake
  • Build real-time and batch data ingestion pipelines from diverse operational systems using high-performance Kafka data pipelines.
  • Implement data transformations that serve digital twin platforms and operational analytics
  • Integrate Kafka event streams with Databricks for real-time operational state updates
  • Implement data quality checks using Delta Live Tables expectations
  • Ensure data governance compliance through Unity Catalog (lineage, access control, metadata)
  • Optimize pipeline performance, reliability, and cost efficiency
  • Write clean, well-documented, and testable code following engineering best practices
  • Collaborate with ML engineers to deliver feature-engineered datasets
  • Participate in code reviews, knowledge sharing, and continuous improvement initiatives
  • Support production data systems through monitoring, troubleshooting, and incident resolution.
  • Build business data warehouse solutions using Terradata for business intelligence.

 

Our core data platform stack includes:

Data Platform & Lakehouse

  • Databricks as the single point of truth for all data
  • Realtime Data Pipelines implemented using Kafka for data ingestion.
  • Databricks SQL for analytical queries
  • Unity Catalog for metadata management and governance
  • Terradata for data warehouse and business intelligence.

Stream & Event Processing

  • Apache Kafka for real-time event ingestion
  • Structured Streaming for continuous data processing
  • Delta Live Tables for declarative, quality-enforced pipelines

Data Quality

  • Delta Live Tables expectations for data validation
  • Data profiling and anomaly detection

Preferred Qualifications

  • 7+ years of hands-on data engineering experience
  • Track record of building and maintaining production-grade data pipelines
  • Experience with Delta Live Tables for declarative pipeline development
  • Experience working in agile, cross-functional teams
  • Familiarity with time-series data patterns and operational data modelling

Highly Desirable

  • Experience building data pipelines for digital twin or simulation platforms
  • Familiarity with operational state modeling for real-time systems
  • Exposure to physics-informed or time-series ML feature engineering
  • Experience working with distributed, multidisciplinary teams
  • Exposure to industrial domains such as Manufacturing, Logistics, or Transportation is a plus