Senior Data Engineer
AgencyBloc Cedar Falls, Iowa, United States
Software Development · 51-200 employees
About the role
The Senior Data Engineer will design, build, and maintain scalable data pipelines on Databricks and AWS while implementing Medallion architecture standards. They will also mentor team members, ensure data quality and governance, and optimize data models to support BI and AI workloads.
What they look for
Requirements
Candidates must have 7+ years of data engineering experience, including at least 2 years at a senior level. Proficiency in Databricks, PySpark, AWS, and dimensional data modeling is required, along with a bachelor's degree in Computer Science or equivalent experience.
Full description
AgencyBloc is a leading provider of agency management solutions built specifically for life and health insurance agencies. Our platform helps agencies streamline operations, strengthen client relationships, and drive growth through powerful CRM, commission processing, and marketing automation tools. We value innovation, collaboration, and making a meaningful impact in the industries we serve and we take pride in fostering a culture that genuinely cares about our customers, our work, and each other.
Requirements
Summary
As a Senior Data Engineer, you will be the senior hands-on builder of AgencyBloc's cloud data platform, implementing the architecture, standards, and patterns defined by the Data Architect and turning them into reliable, production-grade data products.
You will design and build the pipelines, data models, and frameworks that power our Databricks lakehouse on AWS, moving data from our SaaS applications and operational systems (transactional databases, CRM, support, and other internal sources) through bronze, silver, and gold layers to serve BI, analytics, and AI workloads.
This role operates at a senior implementation level: you own delivery of complex data initiatives end-to-end within the established architecture, raise the engineering bar through code review and mentorship of Data Engineers and Data Developers, and contribute practical, build-informed feedback that shapes the platform's standards and roadmap.
Responsibilities:
- Design, build, and maintain scalable batch and incremental data pipelines on Databricks and AWS, following the architectural standards and patterns established by the Data Architect.
- Implement and extend the Medallion (bronze/silver/gold) architecture, including ingestion into raw layers, transformation into refined and conformed layers, and promotion into curated, consumption-ready data products.
- Build and optimize data models (dimensional and other established patterns) in the silver and gold layers to serve BI, reporting, and AI use cases.
- Develop ingestion for new data sources, including SaaS application APIs, CRM and support systems, and relational/OLTP databases, handling schema evolution, incremental loads, and late-arriving or malformed data safely.
- Write clean, tested, well-documented Python/PySpark and SQL, and package reusable frameworks and utilities that make the next pipeline faster to build than the last.
- Own delivery of assigned data initiatives end-to-end: design within established patterns, build, test, deploy, operate, and iterate.
- Contribute improvements to platform standards, tooling, and reusable patterns based on hands-on implementation experience, partnering with the Data Architect.
Data Governance, Quality & Compliance:
- Implement data quality validation, reconciliation, and remediation within pipelines according to platform quality standards, ensuring issues are caught before they reach consumers.
- Apply governance standards in day-to-day work, including cataloging, lineage, classification, documentation, and ownership metadata for the datasets you build.
- Implement secure data handling in accordance with security and compliance requirements, including access controls, encryption, masking, and PII handling for sensitive and regulated data (SOC 2 and related standards).
- Apply data retention, archival, and lifecycle rules to the datasets and pipelines you own.
Observability & Operational Excellence
- Instrument pipelines with monitoring, alerting, and logging aligned to platform observability standards, including SLAs/SLOs for freshness, completeness, and reliability.
- Triage, troubleshoot, and resolve pipeline failures and data incidents, performing root-cause analysis and implementing preventive fixes.
- Build with cost in mind: right-size compute, tune Spark workloads, and monitor the cost profile of the pipelines you own, flagging optimization opportunities.
Platform Standardization & Leadership
- Mentor Data Engineers and Data Developers through code review, pairing, and technical guidance, raising the team's engineering standards.
- Serve as a technical escalation point for complex pipeline, modeling, and performance problems.
- Participate in architecture reviews, providing implementation-level input on feasibility, effort, and operational impact.
- Partner with analytics, product, and engineering stakeholders to understand data needs and translate them into well-modeled, trustworthy datasets.
AI & Analytics Enablement
- Build and maintain the datasets, features, and serving structures that power BI dashboards and AI/ML workloads on the lakehouse.
- Implement patterns to deliver trusted, well-governed data to AI and analytics use cases, as defined in the platform strategy.
Skills/Education/Experience:
- Bachelor's degree in Computer Science or equivalent experience preferred.
- 7+ years of experience in data engineering or analytics engineering, with at least 2 years operating at a senior level (leading delivery of significant data initiatives).
- Strong hands-on production experience with Databricks and/or Spark, including PySpark development, Delta Lake, and workload tuning; Unity Catalog experience is a plus.
- Experience building pipelines on AWS (e.g., S3, IAM, RDS, Glue or comparable services); dual-cloud exposure (Azure) is a plus.
- Experience implementing Medallion (bronze/silver/gold) or comparable layered lakehouse architectures.
- Strong data modeling skills for analytical use cases (dimensional modeling required; exposure to Data Vault or similar history-preserving patterns is a plus).
- Experience extracting and modeling data from relational/OLTP source systems (e.g., MySQL), including CDC or incremental load patterns and handling imperfect source data.
- Expert-level SQL and strong Python, with software engineering fundamentals: version control, testing, code review, and CI/CD for data pipelines (e.g., GitHub Actions).
- Experience implementing data quality checks, pipeline observability, and monitoring/alerting in production.
- Working knowledge of data security practices, including access management, encryption, masking, and PII handling in regulated environments (SOC 2 or similar).
- Experience with orchestration tools (e.g., Databricks Workflows, Airflow, or comparable).
- Strong communication skills, including the ability to explain technical decisions and data issues to non-technical stakeholders.
- Experience working in insurance, InsurTech, or other regulated industries is a plus.
- Experience supporting AI/ML data needs (feature datasets, training data preparation) is a plus.
Applicants for employment in the US must have work authorization that does not now or in the future require sponsorship of a visa for employment authorization in the United States.
Similar roles
- Data Engineer
-
Senior Data Engineer
Datavail Mumbai, Maharashtra, India
-
Data Engineer (m/w/d) - Croatia
AT GmbH City of Zagreb, Croatia
-
Jr Data Engineer
ECS Tech Inc Fairfax, Virginia, United States · $105K–$126K/yr
-
Data Engineer (m/w/d) Geodatenmanagement
aconium GmbH Berlin, Berlin, Germany
-
Data Engineer Business Intelligence (m/w/d)
aconium GmbH Berlin, Berlin, Germany