RADcube - A NLogix Company

Data Engineer (Databricks/AWS)

RADcube - A NLogix Company · Indianapolis, Indiana, United States

IT System Custom Software Development · 201-500 employees

2 h ago
Mid (2-5 yrs) Full-time United States
Log in to apply, save this posting, or score it against your profile with AI.

About the role

Design, build, and maintain scalable ETL/ELT pipelines using Databricks, AWS, and Apache Spark to support analytics and data science initiatives. Collaborate with cross-functional teams to ensure data quality, governance, and performance optimization across production environments.

What they look for

Databricks AWS Apache Spark PySpark ETL/ELT Pipelines Python SQL Big Data Data Orchestration Apache Airflow CI/CD Data Governance Delta Lake Snowflake Infrastructure-as-Code Docker

Requirements

Requires 3–5 years of data engineering experience within the pharma industry and heavy hands-on expertise with Databricks and AWS data services. Candidates must possess strong SQL and Python skills, along with a bachelor's degree in computer science or a related field.

Full description

Data Engineer

Hybrid – Indianapolis, IN

About the Role

We are seeking a Data Engineer with 3–5 years of experience working specifically within the pharma industry to join a pharma-focused data team. This role is responsible for designing, building, and maintaining the ETL/ELT pipelines and data infrastructure that power analytics, reporting, and data science work across the business. You will work closely with data analysts, data scientists, and business stakeholders to ensure data is reliable, well-governed, and readily available for downstream use.

Key Responsibilities

  • Design, build, and maintain scalable ETL/ELT pipelines using Databricks, AWS data services, and Apache Spark
  • Ingest, transform, and load large-scale (Big Data) datasets from a variety of source systems
  • Build and manage data orchestration workflows, including scheduling, monitoring, and failure recovery
  • Implement CI/CD practices for data pipeline development and deployment
  • Ensure data quality, consistency, and governance across pipelines, including validation and schema checks
  • Optimize pipeline performance through partitioning, compression, caching, and tuning strategies
  • Collaborate with data analysts and data scientists to deliver analytics-ready and model-ready datasets
  • Apply pharma domain knowledge to ensure data models and pipelines reflect real business needs
  • Support production pipelines, troubleshoot issues, and perform root cause analysis

Requirements

Required Qualifications

  • 3–5 years of data engineering experience specifically within the pharma industry
  • Heavy hands-on experience with Databricks
  • Heavy hands-on experience with AWS data services (e.g., S3, Glue, Redshift, Lambda, Kinesis)
  • Strong experience with Big Data technologies and Apache Spark (PySpark preferred)
  • Demonstrated experience building and maintaining ETL/ELT pipelines end-to-end
  • Experience with orchestration tools (e.g., Apache Airflow) and CI/CD practices
  • Experience with monitoring/observability for data pipelines
  • Demonstrated ability to understand pharma business needs and speak to pharma business groups
  • Strong SQL and Python skills
  • Bachelor's degree in computer science, Data Engineering, or a related field, or equivalent practical experience

Preferred Qualifications

  • Databricks Data Engineer certification
  • Experience with Delta Lake, Snowflake, or similar modern data platforms
  • Experience with Infrastructure-as-Code and containerization (Docker, Kubernetes)
  • Prior experience supporting pharma commercial, clinical, or R&D data functions