Senior Data Engineer
Procter & Gamble Hyderabad, Telangana, India
Manufacturing · 10,001+ employees
About the role
Architect AI-native data pipelines and build a semantic data layer to support knowledge and cognitive agents. Lead the full application lifecycle for data components while enforcing engineering standards like CI/CD and data quality.
What they look for
Requirements
Requires a Bachelor's degree in a technical field and at least 5 years of experience developing Big Data solutions with Spark and Databricks. Candidates must have hands-on experience with vector databases, RAG pipelines, and modern data architecture.
Benefits
Full description
Job Location
HYDERABAD OFFICE INDIA PSC PGH
Job Description
P&G was founded over 180 years ago as a simple soap and candle company. Today, we're the world’s largest consumer goods company and home to iconic, trusted brands that make life a little bit easier in small but meaningful ways. We've spanned three centuries thanks to three simple ideas: leadership, innovation and citizenship. The insight, innovation and passion of hardworking teams have helped us grow into a global company that is governed responsibly and ethically, that is open and visible, and that supports good causes and protects the environment. This is a place where you can be proud to work and do something that matters.
Dedication from Us: You'll be at the core of breakthrough innovations, be given exciting assignments, lead initiatives, and take ownership and responsibility, in creative workspaces where new insights thrive. All the while, you'll receive outstanding training to help you become a leader in your field. It is not just about what you'll do, but how you'll feel: encouraged, valued, purposeful, challenged, heard, and inspired.
What we Offer: Continuous mentorship – you will collaborate with peers and receive both formal training as well as day-to-day mentoring from your manager dynamic and encouraging work environment– employees are at the centre, we value every individual and support initiative, promoting agility and work/life balance.
The Opportunity We’re looking for a Senior Data Engineer to join our Supply Chain Innovation team and build the trusted data foundation for an AI-native operating model where people and AI work together to improve decisions, execution, and business outcomes. This role goes beyond moving and transforming data for dashboards and reports: you will architect pipelines that feed knowledge, action, and cognitive agents with trusted, realtime, and governed data, and design the semantic layer, KPI logic, and data acquisition strategy as first-class engineering concerns rather than afterthoughts. You will leverage modern Agile and DevOps practices to lead the design and development of AI-native data systems, delivering projects in multinational teams.
Position Responsibilities
- Architect AI-native data pipelines — design end-to-end ingestion, transformation, and serving pipelines optimized for both traditional analytics consumption and agentic consumption (retrieval-augmented generation, tool-calling, MCP-based access).
- Build the semantic data layer — standardize and contextualize data (schemas, metadata, business definitions) so knowledge and cognitive agents can reliably interpret and reason over it.
- Own the KPI factory / manipulation layer — encode standard KPI calculations and business logic in reusable back-end services so agents and downstream consumers retrieve trusted outputs instead of recalculating in real time.
- Design the data acquisition strategy — define systems of record, reference data, access rules, refresh frequency, and ownership for every data element feeding AI agents.
- Build agent retrieval infrastructure — implement vector stores, embedding pipelines, and RAG patterns that support knowledge agents at production scale.
- Govern agent data access — Design permissioning, guardrails, approval workflows, and audit logging for AI-assisted and automated data interactions, ensuring appropriate human accountability for sensitive or consequential actions.
- Lead the full application lifecycle — from development, deployment, and upgrade through replacement or termination, for AI-native data pipeline components on Azure.
- Enforce engineering standards — CI/CD, testing, data quality checks, drift detection, and evaluation of agent-facing outputs, in line with modern DevOps and Agile practices.
- Collaborate cross-functionally — partner with agentic engineers, platform engineers, governance and risk leads, and business process experts to translate work-process redesign into production-grade data architecture.
- Design for Human + AI Collaboration - Design data products and access patterns that improve how people work with AI—reducing repetitive data preparation and retrieval while preserving human judgment, accountability, and control for decisions and actions that require it.
The Ideal Candidate
Most data engineering roles optimize reporting and BI consumption. This one is built for a world where people and AI-enabled applications work together as consumers of the data platform. That means designing for real-time trust, explainability, and governed access, so that knowledge agents can retrieve accurate context, action agents can execute safely, and cognitive agents can reason over data that is current, standardized, and auditable with human in the loop.
Job Qualifications
Qualifications:
- Bachelor’s Degree with a related technical field, or equivalent practical experience.
- At least 5+ years of experience developing Big Data solutions using Spark and Databricks, including recent experience designing data products or pipelines for AI, ML, decision-support, or AI-assisted workflows.
- Extensive knowledge of modern Big Data architecture: Data Warehousing, Lakehouse, Data Mesh, Data Quality, and semantic layer design.
- Hands-on experience with vector databases and embeddings (e.g. Azure AI Search, pgvector, FAISS, or equivalent) and RAG pipeline design.
- Fluency in Python (PySpark) and SQL, plus working knowledge of LLM orchestration frameworks (e.g. LangChain, LangGraph , LlamaIndex, or equivalent) and context engineering for data-facing agents.
- Working knowledge of Airflow and Azure Data Factory, with experience integrating agent orchestration and tooling layers (e.g. MCP servers, tool-calling APIs).
- Experience designing solutions including unit testing, CI/CD pipelines, and API integrations, with added rigor around data lineage, auditability, and governance for AI-consumed data.
- Understanding of AI-native operating model concepts — in the loop, on the loop, and governing the loop — and how they translate into data access and approval architecture.
- Strong verbal, written, and interpersonal communication skills, with the ability to work across engineering, governance, and business process teams.
- A strong desire to produce high-quality, trustworthy data infrastructure through cross-functional collaboration, testing, code reviews, and other best practices.
About Us
We produce globally recognized brands, and we grow the best business leaders in the industry. With a portfolio of trusted brands as diverse as ours, it is paramount our leaders can lead with courage the vast array of brands, categories, and functions. We serve consumers around the world with one of the strongest portfolios of trusted, quality, leadership brands, including Always®, Ariel®, Gillette®, Head & Shoulders®, Herbal Essences®, Oral-B®, Pampers®, Pantene®, Tampax® and more. Our community includes operations in approximately 70 countries worldwide.
Visit http://www.pg.com to know more.
We are an equal-opportunity employer and value diversity at our company. We do not discriminate against individuals based on race, color, gender, age, national origin, religion, sexual orientation, gender identity or expression, marital status, citizenship, disability, HIV/AIDS status, or any other legally protected factor.
At P&G, the hiring journey is personalized every step of the way, thereby ensuring equal opportunities for all, with a strong foundation of Ethics & Corporate Responsibility guiding everything we do.
All the available job opportunities are posted either on our website - pgcareers.com, or on our official social media pages, for the convenience of prospective candidates, and do not require them to pay any kind of fees towards their application.”
Job Schedule
Full time
Job Number
R000158486
Job Segmentation
Experienced Professionals
Similar roles
-
Senior AI-Native Data Engineer (f/m/x)
exmox Hamburg, Germany
-
Data Engineer
Ford Motor Company Chennai, Tamil Nadu, India
-
Cloud Data Engineer
EXL Gurugram, Haryana, India
-
Data Engineer - DBT & SQL
Zensar Pune, Maharashtra, India
-
Data Engineer (REF5635Y)
Deutsche Telekom IT Solutions Budapest, Central Hungary, Hungary
-
Data Engineer
Entain Gibraltar, Gibraltar