Lead Data Engineer
ClinDCast LLC · Warsaw, Indiana, United States · $114K–$125K/yr
IT Services and IT Consulting · 51-200 employees
About the role
The Lead Data Engineer will design and build scalable data platforms, pipelines, and architectures for enterprise-scale solutions. They are responsible for ensuring data quality, managing pipeline operations, and implementing DevOps practices to improve system reliability.
What they look for
Requirements
Candidates must have extensive experience in ETL/ELT development, cloud-based data engineering, and modern data warehouse architectures. Proficiency in Python, Scala, or Java and strong knowledge of data modeling and streaming data processing are required.
Full description
Job Summary
We are seeking an experienced Data Engineering Architect / Lead Data Engineer to design, build, and optimize scalable data platforms and pipelines. The ideal candidate will have extensive experience in ETL/ELT development, cloud-based data engineering, real-time and batch data processing, and modern data warehouse architectures. This role requires strong expertise in designing end-to-end data solutions that are scalable, secure, and aligned with enterprise standards.
Required Skills & Qualifications
- Advanced expertise in designing and developing ETL/ELT pipelines.
- Strong experience with batch data processing and near real-time/streaming data pipelines.
- Hands-on experience working with structured and semi-structured data.
- Strong knowledge of:
- Incremental data loading
- Change Data Capture (CDC)
- Pipeline orchestration and dependency management
- Strong programming skills in Python (preferred), Scala, or Java.
- Experience optimizing large-scale data processing workloads for performance and cost.
- Solid understanding of data modeling concepts:
- Star Schema
- Snowflake Schema
- Normalized and denormalized data models
- Hands-on experience with at least one major cloud platform:
- Microsoft Azure
- Amazon Web Services (AWS)
- Google Cloud Platform (GCP)
- Strong experience with modern data warehouses such as:
- Snowflake
- Azure Synapse
- Google BigQuery
- Amazon Redshift
Key Responsibilities
Data Architecture & Solution Design
- Design end-to-end data engineering architectures for enterprise-scale solutions.
- Develop scalable architectures for:
- Data Lakes and Lakehouse platforms
- Enterprise Data Warehouses
- Streaming and real-time data processing systems
- Ensure solutions align with enterprise architecture, security, governance, and compliance standards.
- Review and approve technical designs and implementation strategies.
Data Pipeline Development & Management
- Lead the design and development of scalable ETL/ELT pipelines.
- Build and manage data ingestion pipelines for both batch and real-time data.
- Process structured and semi-structured data efficiently.
- Optimize data pipelines for performance, reliability, scalability, and cost.
- Manage schema evolution, metadata, and pipeline dependencies.
Data Quality, Reliability & Operations
- Establish and enforce data quality standards and validation frameworks.
- Implement monitoring, alerting, logging, and observability for data pipelines.
- Perform root cause analysis and resolve data-related production issues.
- Drive operational excellence by improving system stability and reliability.
DevOps / DataOps
- Build and maintain CI/CD pipelines for data engineering workloads.
- Automate testing, deployment, and rollback processes.
- Improve platform reliability and deployment efficiency through automation and DevOps best practices.
Preferred Qualifications
- Experience with modern DataOps and CI/CD practices.
- Knowledge of data governance, security, and compliance frameworks.
- Experience designing enterprise-scale cloud-native data platforms.
- Strong analytical, troubleshooting, and communication skills.