Amgen

Associate Data Engineer

Amgen Hyderabad, Telangana, India

Biotechnology Research · 10,001+ employees

12 h ago
data-engineer Mid (2-5 yrs) Full-time India
Create a free account to apply — email only, no card. You can also save this posting or score it against your profile with AI.

About the role

The role involves designing, building, and maintaining scalable data pipelines and ETL processes to ensure data quality and accessibility. You will collaborate with cross-functional teams to develop data models and implement security measures to support business decision-making.

What they look for

Databricks Python PySpark Scala SQL ETL Data Pipelines AWS PostgreSQL MySQL Apache Hadoop Spark Kafka Data Modeling Data Warehousing Data Integration

Requirements

Candidates must hold a bachelor's degree in Computer Science or a related field with 2 to 4 years of relevant experience. Proficiency in Databricks, Python, PySpark, Scala, and SQL is required, along with a strong understanding of cloud platforms like AWS.

Full description

Career Category

Information Systems

Job Description

Role Description:

The role is responsible for designing, building, maintaining, analyzing, and interpreting data to provide actionable insights that drive business decisions. This role involves working with large datasets, developing reports, supporting and executing data governance initiatives and, visualizing data to ensure data is accessible, reliable, and efficiently managed. The ideal candidate has strong technical skills, experience with big data technologies, and a deep understanding of data architecture and ETL processes

Roles & Responsibilities:

  • Design, develop, and maintain data solutions for data generation, collection, and processing
  • Be a key team member that assists in design and development of the data pipeline
  • Create data pipelines and ensure data quality by implementing ETL processes to migrate and deploy data across systems
  • Contribute to the design, development, and implementation of data pipelines, ETL/ELT processes, and data integration solutions
  • Take ownership of data pipeline projects from inception to deployment, manage scope, timelines, and risks
  • Collaborate with cross-functional teams to understand data requirements and design solutions that meet business needs
  • Develop and maintain data models, data dictionaries, and other documentation to ensure data accuracy and consistency
  • Implement data security and privacy measures to protect sensitive data
  • Very good understanding of Databricks
  • Leverage cloud platforms (AWS preferred) to build scalable and efficient data solutions
  • Collaborate and communicate effectively with product teams
  • Collaborate with Data Architects, Business SMEs, and Data Scientists to design and develop end-to-end data pipelines to meet fast paced business needs across geographic regions
  • Identify and resolve complex data-related challenges
  • Adhere to best practices for coding, testing, and designing reusable code/component
  • Explore new tools and technologies that will help to improve ETL platform performance
  • Participate in sprint planning meetings and provide estimations on technical implementation

Basic Qualifications and Experience:

  • Bachelor’s degree and 2 to 4 years of Computer Science, IT or related field experience

Functional Skills:

Must-Have Skills

  • Proficiency in Databricks ,Python, PySpark, and Scala for data processing and ETL (Extract, Transform, Load) workflows, with hands-on experience in using Databricks for building ETL pipelines and handling big data processing
  • Strong knowledge of SQL and experience with relational (e.g., PostgreSQL, MySQL) databases.
  • Familiarity with big data frameworks like Apache Hadoop, Spark, and Kafka for handling large datasets.

Good-to-Have Skills:

  • Experience with cloud platforms such as AWS particularly in data services (e.g., EKS, EC2, S3, EMR, RDS, Redshift/Spectrum, Lambda, Glue, Athena)
  • Understanding of data modeling, data warehousing, and data integration concepts
  • Understanding of machine learning pipelines and frameworks for ML/AI models

.

Similar roles