DATAECONOMY

Data Engineer — ML Training Data Pipeline

DATAECONOMY Hyderabad, Telangana, India

Information Technology & Services · 201-500 employees

13 h ago
data-engineer Senior (5-10 yrs) Full-time India
Create a free account to apply — email only, no card. You can also save this posting or score it against your profile with AI.

About the role

Build and maintain end-to-end data pipelines that transform raw production traces into high-quality training datasets for LLM fine-tuning. Responsibilities include implementing data ingestion, deduplication, quality gating, and identity-aware train/test splitting at scale on AWS.

What they look for

Python Pandas Pyarrow JSONL AWS S3 AWS EC2 Data engineering ML data pipelines HuggingFace Datasets Arrow-based storage Data deduplication Data validation LLM fine-tuning Tool-calling schemas Chat templates

Requirements

Requires 6+ years of experience in data engineering with a strong focus on ML data pipelines and Python. Proficiency in AWS services, large-scale JSONL processing, and ML data libraries like HuggingFace Datasets is essential.

Benefits

Comprehensive medical coverage Group personal accident insurance Group term life insurance Retirement benefits Flexible work options Generous leave policy Employee well-being spaces

Full description

Job Title:Data Engineer - ML Training Data Pipeline

Notice period: 0-30 Days

Experience : 5+ Years

Location: Hyderabad OR Pune

We are looking for Data Engineer - ML Training Data Pipeline who can Build and maintain the data pipeline that transforms raw production traces into high-quality training datasets for LLM fine-tuning-ingestion, deduplication, format conversion, quality filtering, and train/test splitting at scale on AWS.

What We Expect:

  • Build

end-to-end data pipelines: raw trace ingestion → dedup → format conversion → quality gating → training-ready datasets

  • Process

large-scale JSONL data on AWS S3 (tens of thousands of traces per batch)

  • Convert

between chat-completion formats (e.g., OpenAI → Llama 3.1 tool-calling format)

  • Implement

smart deduplication and sampling to balance training distribution

  • Design

identity-aware train/test splits that measure true generalization

  • Build

data validation gates to detect schema drift and format anomalies

  • Create

a continuous pipeline that auto-processes new production traces for retraining

Requirements

  • Experience: 6+

years data engineering focused on ML data pipelines

  • Python: Strong — pandas,

pyarrow, JSONL processing at scale

  • ML

Data Libraries: HuggingFace Datasets, Arrow-based storage

  • Data

Formats: Multi-turn conversation/chat data structures and tokenizer-specific formatting

  • Deduplication: Content

hashing, identity-based grouping strategies

  • AWS: S3,

EC2, batch processing workflows

Preferred (Not Required): LLM training data prep (chat templates, tool-calling schemas); Axolotl or similar dataset formats; data versioning (DVC, LakeFS); browser-automation trace data or Playwright.

Benefits

  • Comprehensive Medical Coverage:

Health insurance of INR 5.0 Lakhs for you and your family (up to 6 members), ensuring complete peace of mind.

  • Robust Protection Plans:

Group Personal Accident Insurance and Group Term Life Insurance to safeguard you and your loved ones.

  • Retirement Benefits:

PF and Gratuity provided as per standard government regulations.

  • Flexible Work Options:

Enjoy hybrid work arrangements & flexible working hours.

  • Generous Leave Policy:

21 days of annual leave, in addition to 10 company-declared holidays.

  • Employee Well-being Spaces:

Access to a dedicated break-out area with round-the-clock refreshments for relaxation and rejuvenation.

Similar roles