AstraZeneca

Associate Data Engineer

AstraZeneca Chennai, Tamil Nadu, India

Pharmaceutical Manufacturing · 10,001+ employees

20 h ago
data-engineer Senior (5-10 yrs) Full-time India
Log in to apply, save this posting, or score it against your profile with AI.

About the role

You will implement and monitor end-to-end data pipelines and ETL jobs to ensure high-quality data delivery at scale. Additionally, you will manage cloud data storage, perform root cause analysis on pipeline failures, and drive automation to improve reliability.

What they look for

Python Pandas PySpark SQL AWS S3 AWS Redshift AWS EMR Postman ETL Data Engineering API Orchestration Data Quality Troubleshooting Automation DevOps Git

Requirements

The role requires 5-8 years of experience in data engineering or production support with proficiency in Python, SQL, and AWS cloud services. Candidates must possess strong analytical problem-solving abilities and experience with API-based job execution and data governance.

Full description

Job Title: Associate Data Engineer  

GCL : C3  

Introduction to role:  

Are you ready to keep mission-critical data flowing at global scale and turn incidents into improvements that boost reliability? As an Associate Data Engineer, you will run and enhance the cloud pipelines that power decisions across 85+ markets, ensuring timely, high-quality data reaches the people who need it most.

You will join a high-performing, digitally savvy team that partners across the enterprise to drive speed and precision. Your focus on automation, monitoring, and rapid incident response will translate into trusted data and smoother releases—accelerating how we deliver life-changing medicines. Can you picture yourself orchestrating robust pipelines that help colleagues move faster with confidence?  

Accountabilities:  

Pipeline Operations: Implement and monitor end-to-end data pipelines and ETL jobs across multiple stages to ensure on-time, high-quality delivery at scale.

Data Transformation: Maintain and modify Python (Pandas, PySpark) scripts in line with evolving business needs to improve data quality and performance.

Cloud Data Management: Manage data storage and protected data exchanges across AWS S3, Redshift, and EMR, keeping data flows accurate and compliant.

API Orchestration: Trigger and validate jobs using Postman and other API interfaces to keep schedules on track and detect issues early.

Data Flow Governance: Track inbound and outbound files, log exceptions, and maintain observability to prevent and detect data breaks.

Incident Response and Root Cause Analysis: Investigate and remediate pipeline failures or delays, implement durable fixes, and drive automation that reduces repeat incidents.

Teamwork and Collaborator Management: Work closely with data providers, data custodians, and DevOps teams to assure pipeline health and data accuracy across global collaborators.

Documentation and Versioning: Keep pipeline documentation, job schedules, and technical configurations up to date; support code enhancements and environment updates.

Quality Control: Participate in data quality procedures with Data Stewards to validate releases and safeguard trust in data products.

Continuous Improvement: Identify and implement opportunities to standardize, simplify, and automate operations, increasing reliability and throughput over time.  

Essential Skills/Experience:  

Python (PyCharm, Pandas, PySpark) for maintaining ETL scripts and automation routines  

Postman for testing and triggering API-based job executions  

SQL proficiency using tools such as DBeaver to query and validate relational data

AWS services proficiency across Redshift, S3, and EMR for processing and storage  

WinSCP or equivalent tools for secure file transfers  

Proactive, structured approach to monitoring and troubleshooting

Strong programming and analytical problem-solving abilities

Excellent documentation and organizational skills

Ability to work independently and coordinate across functional teams

Desirable Skills/Experience:  

Familiarity with Git and version control systems  

5–8 years of experience in data engineering, production support, or data operations  

Background handling large-scale data workflows in cloud environments  

Experience working in pharmaceutical or healthcare data ecosystems

Consistent track record resolving performance bottlenecks and job failures

Familiarity with DevOps principles and agile ways of working  

Why AstraZeneca:  

Here, data engineering fuels real-world impact. You’ll work with modern cloud platforms and digital tools, side by side with unexpected combinations of experts—engineers, data stewards, and market teams in the same room—turning bold ideas into operational reality. We move with urgency and clarity, blending imagination with rigor to strengthen how the business runs today while preparing for tomorrow. Your contribution will help colleagues across the globe focus on what matters most, translating into faster, smarter decisions that ultimately benefit patients. We value patience alongside ambition, and we back curiosity with the support and autonomy needed to deliver significant results.

Call to Action:

If you’re ready to build resilient workflows that drive faster decisions and tangible patient impact, step forward and build what reliable data can make possible!

Date Posted

12-Aug-2026

Closing Date

27-Aug-2026

AstraZeneca embraces diversity and equality of opportunity.  We are committed to building an inclusive and diverse team representing all backgrounds, with as wide a range of perspectives as possible, and harnessing industry-leading skills.  We believe that the more inclusive we are, the better our work will be.  We welcome and consider applications to join our team from all qualified candidates, regardless of their characteristics.  We comply with all applicable laws and regulations on non-discrimination in employment (and recruitment), as well as work authorization and employment eligibility verification requirements.

Similar roles