WARWICK INVESTMENT GROUP LLC

Data Engineer

WARWICK INVESTMENT GROUP LLC · Oklahoma City, Oklahoma, United States

Investment Management · 51-200 employees

20 h ago
Senior (5-10 yrs) Full-time United States
Log in to apply, save this posting, or score it against your profile with AI.

About the role

The Data Engineer will design and maintain scalable data ingestion pipelines and warehouse architectures to support BI reporting and AI-driven decision-making. They will also implement data quality, governance, and AI-ready data structures to ensure a reliable single source of truth.

What they look for

SQL Python Data Modeling Snowflake Azure Airflow dbt Data Engineering Oil and Gas Data Governance CI/CD Docker RAG LLMs Data Pipelines Cloud Infrastructure

Requirements

Candidates must have 5+ years of experience in oil and gas data environments with advanced proficiency in SQL, Python, and cloud data platforms like Snowflake and Azure. Strong expertise in data modeling, pipeline orchestration, and data governance is required to succeed in this role.

Full description

Job DetailsLevel: ExperiencedJob Location: Oklahoma City, OK 73116Position Type: Full TimeJOB TITLE: Data Engineer FIRM INTRODUCTION Warwick Investment Group is a private equity firm focused on investing in real assets. With approximately $1.5 billion in managed assets and a growing team of professionals, Warwick combines industry expertise with innovative technology, data analytics, and machine learning to drive intelligent investment decisions and maximize asset value. Data is embedded in every aspect of our business. We leverage extensive public and private datasets to identify opportunities, optimize asset management strategies, and continuously evaluate risk. This position offers a unique opportunity to work cross-functionally with multiple departments while expanding your knowledge of the oil and gas industry. JOB OVERVIEW The Data Engineer at Warwick Energy owns the infrastructure that moves data from source systems into a trusted, well-modeled warehouse, and prepares that data for consumption by both human analysts and AI systems. You will design and maintain ingestion pipelines, orchestrate ELT workflows, enforce data quality, and build the semantic and structural layers that let BI tools, machine learning models, and AI agents draw on a single source of truth. A critical dimension of this role is forward-looking: as the organization moves toward AI-driven decision-making, you will shape data assets so they are discoverable, well-documented, and structured for retrieval-augmented generation (RAG), agent workflows, and automated analytics. You will partner with data scientists, BI analysts, and business stakeholders to ensure the data platform scales with both traditional reporting needs and emerging AI use cases. KEY JOB RESPONSIBILITIES

  • Pipeline Development and Orchestration: Build, monitor, and maintain data ingestion pipelines from source systems (APIs, databases, flat files, SCADA/IoT) into the data warehouse. Orchestrate ELT workflows using tools like Coalesce, dbt, Airflow, or Prefect, with version control and CI/CD practices.
  • Data Modeling and Warehouse Architecture: Design scalable dimensional and normalized data models. Own the warehouse layer structure (raw, staging, marts) and ensure models support both BI reporting and AI/ML consumption patterns.
  • Data Quality and Governance: Implement data quality checks, monitoring, and alerting across pipelines. Enforce governance standards including lineage tracking, access control, and documentation to maintain trust in data assets.
  • AI-Ready Data Architecture: Structure and document data assets so they are consumable by LLMs, RAG pipelines, and AI agents. Design metadata layers, semantic descriptions, and context-rich schemas that allow AI systems to discover and reason over organizational data.
  • AI-Accelerated Engineering: Use AI coding tools (Claude Code, Copilot) and agent workflows to accelerate pipeline development, automate documentation, generate and validate SQL transformations, and build MCP servers or similar interfaces that expose data to AI systems.
  • Infrastructure and Platform Reliability: Manage cloud data infrastructure (Snowflake, Azure). Monitor pipeline health, optimize query performance, and maintain SLAs for data freshness and availability.
  • Documentation and Collaboration: Produce clear technical documentation for pipelines, data models, and ELT processes. Partner with BI analysts, data scientists, and business teams to align data infrastructure with analytical and AI-driven objectives.

QualificationsREQUIRED SKILLS

  • Oil and Gas Experience: 5+ years in oil and gas data environments, with familiarity across production, land, SCADA, and well data domains.
  • SQL and Data Modeling: Advanced SQL proficiency (CTEs, window functions, query optimization). Experience with dimensional and normalized modeling approaches.
  • Pipeline and Orchestration: Experience building and maintaining ELT/ETL pipelines with Coalesce, dbt, Airflow, Prefect, or similar tools.
  • Python: Proficient in Python for data manipulation, pipeline scripting, and automation tasks.
  • Cloud Data Platforms: Experience with Snowflake & Azure for warehousing, storage, and compute.
  • Version Control and CI/CD: Familiarity with Git, Azure DevOps, or similar systems. Experience with CI/CD for data pipeline deployments.
  • Data Governance: Understanding of data quality frameworks, lineage tracking, access control, and compliance requirements.
  • Documentation: Capable of producing clear technical documentation for pipelines, schemas, and processes for both technical and non-technical audiences.
  • Containerization: Proficiency with Docker for containerizing data workloads and pipeline components, with familiarity deploying containers to cloud services.

DESIRED SKILLS

  • AI-Assisted Development: Experience using AI coding tools (e.g., Claude Code, GitHub Copilot) to accelerate SQL development, dbt workflow creation, and data pipeline work.
  • AI Data Architecture: Familiarity with RAG patterns, vector stores, embeddings, or building data interfaces for LLM and agent consumption.
  • Agent and MCP Workflows: Experience building or orchestrating multi-agent AI systems, MCP servers, or tool-use interfaces that expose data to AI systems.
  • AI-Powered Documentation: Using AI tools to auto-generate and maintain data dictionaries, lineage documentation, and schema descriptions that stay current with the codebase.
  • BI Tool Familiarity: Working knowledge of Power BI or Spotfire to collaborate effectively with the BI team.
  • Streaming and CDC: Experience with change data capture, event streaming (Kafka, Azure Event Hubs), or real-time data pipelines.