Data Engineer (Databricks/AWS)
RADcube - A NLogix Company · Indianapolis, Indiana, United States
IT System Custom Software Development · 201-500 employees
About the role
Design, build, and maintain scalable ETL/ELT pipelines using Databricks, AWS, and Apache Spark to support analytics and data science initiatives. Collaborate with cross-functional teams to ensure data quality, governance, and performance optimization across production environments.
What they look for
Requirements
Requires 3–5 years of data engineering experience within the pharma industry and heavy hands-on expertise with Databricks and AWS data services. Candidates must possess strong SQL and Python skills, along with a bachelor's degree in computer science or a related field.
Full description
Data Engineer
Hybrid – Indianapolis, IN
About the Role
We are seeking a Data Engineer with 3–5 years of experience working specifically within the pharma industry to join a pharma-focused data team. This role is responsible for designing, building, and maintaining the ETL/ELT pipelines and data infrastructure that power analytics, reporting, and data science work across the business. You will work closely with data analysts, data scientists, and business stakeholders to ensure data is reliable, well-governed, and readily available for downstream use.
Key Responsibilities
- Design, build, and maintain scalable ETL/ELT pipelines using Databricks, AWS data services, and Apache Spark
- Ingest, transform, and load large-scale (Big Data) datasets from a variety of source systems
- Build and manage data orchestration workflows, including scheduling, monitoring, and failure recovery
- Implement CI/CD practices for data pipeline development and deployment
- Ensure data quality, consistency, and governance across pipelines, including validation and schema checks
- Optimize pipeline performance through partitioning, compression, caching, and tuning strategies
- Collaborate with data analysts and data scientists to deliver analytics-ready and model-ready datasets
- Apply pharma domain knowledge to ensure data models and pipelines reflect real business needs
- Support production pipelines, troubleshoot issues, and perform root cause analysis
Requirements
Required Qualifications
- 3–5 years of data engineering experience specifically within the pharma industry
- Heavy hands-on experience with Databricks
- Heavy hands-on experience with AWS data services (e.g., S3, Glue, Redshift, Lambda, Kinesis)
- Strong experience with Big Data technologies and Apache Spark (PySpark preferred)
- Demonstrated experience building and maintaining ETL/ELT pipelines end-to-end
- Experience with orchestration tools (e.g., Apache Airflow) and CI/CD practices
- Experience with monitoring/observability for data pipelines
- Demonstrated ability to understand pharma business needs and speak to pharma business groups
- Strong SQL and Python skills
- Bachelor's degree in computer science, Data Engineering, or a related field, or equivalent practical experience
Preferred Qualifications
- Databricks Data Engineer certification
- Experience with Delta Lake, Snowflake, or similar modern data platforms
- Experience with Infrastructure-as-Code and containerization (Docker, Kubernetes)
- Prior experience supporting pharma commercial, clinical, or R&D data functions