Data Engineer
Strategic Innovation Group LLC Arlington County, Virginia, United States
IT Services and IT Consulting · 11-50 employees
About the role
The Senior Data Engineer will design and implement a centralized PostgreSQL data layer and architect scalable ingestion pipelines for a federal analytics platform. They will also set engineering standards, lead data quality efforts, and collaborate with cross-functional teams in an Agile environment.
What they look for
Requirements
Candidates must have at least 8 years of professional data engineering experience and expert-level proficiency in SQL and Python. A bachelor's degree in a technical field is required, along with the ability to obtain a Top Secret clearance.
Benefits
Full description
Must be a U.S. Citizen with an Active TS or eligible for a Top Secret (Tier 5) clearance. No third party.
ABOUT THE ROLE
Strategic Innovation Group (SIG) is seeking a Senior Data Engineer to lead the data core of a new analytics platform for a federal government client. The platform consolidates program, financial, and performance data from multiple federal agency systems and Government-wide sources into a centralized PostgreSQL data layer that feeds dashboards, analytic rules, and reporting for oversight users, and is built on-premises within the client's FedRAMP Moderate environment. This role owns the physical data model, the ingestion architecture, and the reusable connector framework, sets engineering standards for the data team, and is the person who can explain why the numbers on the dashboard are right. The successful candidate will be proposed as key personnel and will work in an Agile environment alongside the program manager, solutions architect, data governance lead, application developers, configuration manager, security lead, and client stakeholders to deliver a Minimum Viable Product and the foundation for its expansion.
ABOUT SIG
SIG is a fast growing 8(a) government contractor based in Arlington, Virginia. We offer a broad range of technical expertise and experience in Digital Transformation, Data Management/Data Science, and Systems Modernization. At SIG, our people are our mission. Come join our team! A successful candidate will be offered the following:
- Great work/life balance - Eligibility for performance-based participation in cash bonuses - Potential to participate in growth of the company through incentives - Excellent benefits, including health, dental, vision, generous PTO, a 401(k) with match, life insurance, short- and long-term disability, and a health savings account (HSA)
RESPONSIBILITIES
- Design and implement the centralized data layer in PostgreSQL to a governed schema and shared data vocabulary defined with the data governance lead. - Architect a scalable ingestion pipeline in Python (Airflow, Dagster, Prefect, or comparable orchestration) that collects, normalizes, validates, and persists structured and unstructured data from agency financial-management, program-management, and payment systems and from Government-wide data sources. - Build configuration-driven connectors (REST APIs, flat files, database replication, SFTP) that are reusable across agencies without source-code changes, and deploy them through the program CI/CD pipeline (GitLab, Jenkins or Azure DevOps) with automated build, test, and security scanning to the client's Rancher-managed Kubernetes platform. - Implement PII anonymization, data-handling, and access controls at ingestion in coordination with the security lead, aligned to NIST SP 800-53 Moderate controls. - Implement defined analytic rules and threshold-based indicators as deterministic, inspectable SQL that runs on a scheduled refresh, and document the data and criteria behind each rule. - Lead the error-detection methodology for data-quality issues: identify, categorize, track, resolve, and report defects and anomalies throughout the lifecycle. - Capture metadata and data lineage, and expose data-currency indicators (source, last refreshed, next refresh) to the user interface. - Set coding, code-review, testing, and documentation standards for the data team; review the Data Engineer's work; and collaborate with application developers on API access to the data core. - Troubleshoot and resolve pipeline failures, data discrepancies, and query performance bottlenecks. - Produce technical documentation, data dictionaries, and runbooks, and deliver knowledge transfer so client staff can operate, maintain, and extend pipelines with minimal contractor reliance. - Participate in Agile ceremonies, sprint planning, code reviews, demonstrations, and continuous improvement activities.
REQUIRED QUALIFICATIONS
- Bachelor's degree in Computer Science, Information Systems, Engineering, or a related field; equivalent experience may be considered. - 8+ years of professional data engineering experience, including 3+ years as technical lead or senior engineer on a production data platform (6+ years with a Master's degree). - Expert-level SQL and PostgreSQL data modeling, including normalized and analytic schemas, partitioning, indexing, and performance tuning. - Strong Python for pipeline development and hands-on experience with an orchestration framework such as Airflow, Dagster, or Prefect. - Demonstrated experience designing reusable, configuration-driven ingestion from heterogeneous sources with schema validation and data-quality checks. - Experience deploying pipelines in containerized environments (Docker, Kubernetes) with Git-based version control and CI/CD. - Experience handling PII and sensitive financial or program data under NIST SP 800-53 or equivalent controls. - Experience working in Agile software development environments and collaborating with technical and non-technical stakeholders in a federal contracting environment. - Strong written communication skills for data-model documentation, runbooks, and knowledge transfer.
PREFERRED QUALIFICATIONS
- Experience supporting federal civilian, financial-oversight, budget, or program-management IT systems. - Experience integrating data from enterprise financial systems (e.g., SAP S/4HANA, Oracle, or comparable ERP) and Treasury or payment platforms. - Experience with grants management or federal financial systems (e.g., GrantSolutions, eRA, Treasury payment systems) highly preferred. - Experience extracting structured data from unstructured documents (PDF, DOCX) at scale. - Experience with metadata management, data-lineage, and data-catalog tooling. - Experience deploying to Kubernetes (Rancher preferred) in an on-premises federal environment with restricted internet access. - Familiarity with Splunk log integration and NIST SP 800-53 audit-logging requirements.
SIG is an Equal Opportunity Employer. All qualified applicants will receive consideration for employment without regard to race, color, religion, sex, sexual orientation, gender identity, national origin, disability, or protected veteran status.
Similar roles
-
Senior Data Engineer
Nordic Investin Group Stockholm, Maine, United States
-
Senior Data Engineer
Okta Bengaluru, Karnataka, India
-
Senior Data Engineer – Apache Spark | Kafka | Flink | Trino | Iceberg | Big Data | Streaming | Data Platform 8–12 Years
Cisco Bengaluru, Karnataka, India
-
Sr Data Engineer, AI
Constellation Energy Generation, LLC. Baltimore, Maryland, United States
-
Data Engineer
Publicis Groupe Santiago, Santiago Metropolitan Region, Chile
-
Senior Data Engineer (USASOC-CDAO)
Kentro Fayetteville, North Carolina, United States