Database Engineer (Remote)
Govcio LLC United States
IT System Data Services · 1,001-5,000 employees
About the role
The Data Engineer will design, build, and maintain enriched datasets using medallion architecture within Databricks to support clinical and operational decision-making. They will also collaborate with stakeholders to gather requirements, manage data pipelines, and ensure high-quality data delivery for Power BI semantic models.
What they look for
Requirements
Candidates must have a bachelor's degree and over 12 years of relevant experience in building data pipelines and transformations. Proficiency in SQL, PySpark, Databricks, and experience with version control tools like GitHub are required.
Full description
GovCIO is currently hiring for Data Engineer to design and build enriched datasets using medallion architecture (bronze/silver/gold) in Databricks, sourcing data from the VA Corporate Data Warehouse (CDW) and other internal and external systems. The datasets you build will feed directly into Power BI semantic models used by stakeholders across the organization. This role works closely with clients to gather requirements, track work in JIRA, and deliver high-quality, well-documented data pipelines. This is a fully remote position.
Responsibilities
Our Data Engineering team supports multiple offices and programs across the Department of Veterans Affairs by building enriched, analytics-ready datasets that power clinical and operational decision-making across the Veterans Health Administration (VHA). We work at the intersection of enterprise data warehousing and modern cloud analytics, turning complex healthcare data into reliable, reusable data products.
What You'll Do
- Design, build, and maintain enriched datasets following medallion architecture (bronze, silver, gold layers) in Databricks
- Develop, orchestrate, and monitor Databricks jobs and pipelines
- Write efficient, well-structured SQL to transform and model healthcare data
- Prepare and optimize gold-layer datasets for consumption by Power BI semantic models
- Manage code and version control through GitHub, following team branching and review practices
- Log, track, and update work items in JIRA
- Partner directly with clients and stakeholders to understand requirements, clarify data definitions, and validate outputs
- Troubleshoot data quality issues and ensure accuracy and consistency across pipelines
- Write ad hoc queries to support client requests, investigations, and one-off analyses
- Perform data validation and quality assurance checks on pipelines and enriched datasets to ensure accuracy, completeness, and consistency
- Document data lineage, transformation logic, and business rules for enriched datasets
- Collaborate with Power BI developers and data scientists on the broader team to ensure datasets meet downstream reporting and analytical needs
Qualifications
Required Skills and Experience
- Bachelors degree and 12+ yrs of experience (or commensurate experience)
- Hands-on experience building data pipelines and transformations in Databricks (PySpark and/or SQL)
- Strong SQL skills, including complex joins, aggregations, and performance tuning
- Experience working with Power BI, including preparing data for semantic models
- Familiarity with GitHub for version control and collaborative development
- Experience using JIRA or similar tools to track and manage work
- Strong communication skills and comfort working directly with clients/stakeholders
- Understanding of medallion (bronze/silver/gold) data architecture principles
Preferred Skills and Experience
- Prior experience working within the VA or VHA
- Deep understanding of the VA Corporate Data Warehouse (CDW)
- Familiarity with VistA/CPRS
- Familiarity with the Federal EHR (Oracle Health)
- Experience with healthcare data domains (e.g., appointments, consults/referrals, orders, TIU notes, visits)
What We're Looking For
Someone who can work independently on complex data problems, communicate clearly with non-technical clients, and take ownership of datasets from raw source through to a polished, analytics-ready product. Healthcare and federal data experience is a strong plus given the sensitivity and complexity of the data we work with.
Posted Salary Range
USD $130,000.00 - USD $160,000.00 /Yr.
Similar roles
-
PostgreSQL Database Administrator
Sectigo Manchester, England, United Kingdom · £50K–£54K/yr
-
Sr. Database Engineer
BANNER SEVENTEEN Boston, Massachusetts, United States · $150K–$175K/yr
-
Azure Infrastructure Database Administrator
CACI $75K–$158K/yr
-
Database Administrator (MariaDB/MySQL/Cassandra/SQL Server)
Support Services Group Hermosillo, Sonora, Mexico
-
Senior Oracle Database Administrator
Ericsson Bengaluru, Karnataka, India
-
Senior Database Administrator - Heart and Vascular
Providence Seattle, Washington, United States · $113K–$179K/yr