Data Engineer — ML Training Data Pipeline
DATAECONOMY Hyderabad, Telangana, India
Information Technology & Services · 201-500 employees
About the role
Build and maintain end-to-end data pipelines that transform raw production traces into high-quality training datasets for LLM fine-tuning. Responsibilities include implementing data ingestion, deduplication, quality gating, and identity-aware train/test splitting at scale on AWS.
What they look for
Requirements
Requires 6+ years of experience in data engineering with a strong focus on ML data pipelines and Python. Proficiency in AWS services, large-scale JSONL processing, and ML data libraries like HuggingFace Datasets is essential.
Benefits
Full description
Job Title:Data Engineer - ML Training Data Pipeline
Notice period: 0-30 Days
Experience : 5+ Years
Location: Hyderabad OR Pune
We are looking for Data Engineer - ML Training Data Pipeline who can Build and maintain the data pipeline that transforms raw production traces into high-quality training datasets for LLM fine-tuning-ingestion, deduplication, format conversion, quality filtering, and train/test splitting at scale on AWS.
What We Expect:
- Build
end-to-end data pipelines: raw trace ingestion → dedup → format conversion → quality gating → training-ready datasets
- Process
large-scale JSONL data on AWS S3 (tens of thousands of traces per batch)
- Convert
between chat-completion formats (e.g., OpenAI → Llama 3.1 tool-calling format)
- Implement
smart deduplication and sampling to balance training distribution
- Design
identity-aware train/test splits that measure true generalization
- Build
data validation gates to detect schema drift and format anomalies
- Create
a continuous pipeline that auto-processes new production traces for retraining
Requirements
- Experience: 6+
years data engineering focused on ML data pipelines
- Python: Strong — pandas,
pyarrow, JSONL processing at scale
- ML
Data Libraries: HuggingFace Datasets, Arrow-based storage
- Data
Formats: Multi-turn conversation/chat data structures and tokenizer-specific formatting
- Deduplication: Content
hashing, identity-based grouping strategies
- AWS: S3,
EC2, batch processing workflows
Preferred (Not Required): LLM training data prep (chat templates, tool-calling schemas); Axolotl or similar dataset formats; data versioning (DVC, LakeFS); browser-automation trace data or Playwright.
Benefits
- Comprehensive Medical Coverage:
Health insurance of INR 5.0 Lakhs for you and your family (up to 6 members), ensuring complete peace of mind.
- Robust Protection Plans:
Group Personal Accident Insurance and Group Term Life Insurance to safeguard you and your loved ones.
- Retirement Benefits:
PF and Gratuity provided as per standard government regulations.
- Flexible Work Options:
Enjoy hybrid work arrangements & flexible working hours.
- Generous Leave Policy:
21 days of annual leave, in addition to 10 company-declared holidays.
- Employee Well-being Spaces:
Access to a dedicated break-out area with round-the-clock refreshments for relaxation and rejuvenation.
Similar roles
-
Senior AI-Native Data Engineer (f/m/x)
exmox Hamburg, Germany
-
Data Engineer
Ford Motor Company Chennai, Tamil Nadu, India
-
Cloud Data Engineer
EXL Gurugram, Haryana, India
-
Data Engineer - DBT & SQL
Zensar Pune, Maharashtra, India
-
Data Engineer (REF5635Y)
Deutsche Telekom IT Solutions Budapest, Central Hungary, Hungary
-
Data Engineer
Entain Gibraltar, Gibraltar