Meeru AI Inc

Senior Staff Data Engineer

Meeru AI Inc Poland

Software Development · 11-50 employees

11 h ago
Remote data-engineer Principal (10+ yrs) Full-time Contractor Poland
Create a free account to apply — email only, no card. You can also save this posting or score it against your profile with AI.

About the role

You will design and build robust data architectures and pipelines to support AI-native financial intelligence systems. The role involves ensuring data integrity, traceability, and performance while collaborating with AI engineers to ground LLM outputs in verified data.

What they look for

Data Engineering SQL Dbt Python Snowflake BigQuery PostgreSQL Airflow Dagster Prefect Data Modeling Change Data Capture Data Architecture Machine Learning LLM Cloud Infrastructure

Requirements

Candidates must have 10+ years of experience in data engineering with a strong background in SQL, dbt, and large-scale data warehouse management. You should possess deep expertise in building idempotent pipelines and a commitment to data quality, lineage, and auditability.

Full description

About Meeru AI

Meeru AI is building an AI-native platform that transforms how finance and accounting teams operate. We connect to enterprise financial systems — ERPs, CRMs, billing platforms, HRIS — and apply machine learning to turn fragmented operational data into grounded, auditable intelligence for CFOs, controllers, and FP&A leaders.

We deploy on customer terms — SaaS multi-tenant, SaaS single-tenant, and on-premises — across AWS, Azure, and GCP. Our customers are Fortune 500 finance teams who require data isolation, auditability, and compliance.

The Role

We are looking for a Senior AI Engineer to build the AI and intelligence layer — and help uphold the discipline that keeps it honest. Our output sits adjacent to externally reported financials, so "sounds plausible" is not good enough: everything the AI produces must be grounded in, and traceable to, verified data.

You'll build the machine-learning models that learn each customer's patterns, the LLM and agentic systems that produce grounded natural-language output, and help uphold the rigor that keeps that output faithful. You own significant pieces of the layer end-to-end, working closely with our Staff AI Engineer and evaluation engineer, and you build per-customer models without leaking the very signal they're meant to detect.

This is a hands-on engineering role. You turn designs into robust production systems, measure quality rigorously, help turn user feedback into durable improvements, and grow toward staff-level technical ownership.

Location & engagement• Offshore — Open to remote globally, working with our distributed engineering team.

  • Flexible engagement — open to full-time, contract, or contract-to-hire, whichever fits you and the engagement best.
  • US partnership — you partner closely with US-based AI leadership, with a few hours of daily time-zone overlap.

What makes this role different• Grounding is a hard requirement — no fabricated or unsupported output, ever. Faithfulness is a gate, not a tuning goal.

  • Clear separation of concerns — ML surfaces and prioritizes; generated text only states what is supported by verified data or confirmed by a human.
  • Per-customer intelligence — models that learn each business's behavior, with strict anti-leakage discipline.
  • Deploys in customer clouds — including managed LLMs in-VPC, so customer data never leaves their environment.
  • 10+ years in data engineering, including significant time at Staff, Principal or Lead level. You have owned data architecture for a whole system, not just individual pipelines.
  • You have designed a canonical or common data model used by many teams. You know how to define shared entities, hierarchies, dimensions and time periods so that many different sources fit into one model, and many consumers can rely on it.
  • You build so that new customers or sources are added by configuration, not by rewriting code. You have designed mapping frameworks or connector strategies that reuse work rather than fork it, and you can explain when to build versus configure.
  • Expert SQL and dbt (or similar) modeling at scale. Window functions, large joins, incremental and change-data-capture (CDC) models, and performance tuning on large datasets are everyday tools for you.
  • Deep warehouse knowledge: Snowflake, BigQuery and/or PostgreSQL. You understand partitioning, internals, and the cost-versus-performance trade-offs of each.
  • Production orchestration with Airflow, Dagster or Prefect. Your pipelines are idempotent and reproducible: the same inputs always produce the same outputs, and re-runs are safe.
  • Lineage, data contracts, data quality and schema evolution are second nature. You have set up lineage capture (e.g. OpenLineage, dbt exposures), quality tests (dbt tests, Great Expectations), and checks that stop a build when numbers don't reconcile.
  • A firm commitment to correctness and traceability. "Every number ties back to its source" is a standard you hold, even under deadline pressure. Our data sits next to companies' reported financials and must stand up to audit.
  • Strong Python for pipelines, tooling and testing.
  • Experience leading and mentoring data engineers, ideally across distributed or offshore teams. You run design reviews, write standards, and raise the bar without becoming the bottleneck.
  • Clear communication with architects, backend and AI engineers, and product. You can explain and defend architecture decisions in writing and in meetings, in English, with US-based colleagues.
  • Willingness to work part of your day overlapping with US hours. Our leadership is US-based.

Nice-to-haves• Financial and ERP source data. Hands-on experience with NetSuite, SAP, Oracle or Workday data, and an understanding of how financial statements are built and reconciled (general ledger, AP, payroll, accruals, prepaids, equity).

  • Multi-cloud and customer-hosted deployments. You have run data workloads across AWS, Azure and GCP, including inside a customer's own cloud with Docker/containers, strict tenant isolation and no data leaving their environment.
  • FinTech or financial-services background, including exposure to SOC 2 or audit expectations for data.
  • Synthetic or "golden" test datasets. You have built realistic test data so teams could develop and regression-test before real customer data was available.
  • Feeding ML or LLM systems from a curated data layer. You understand what AI teams need from clean, well-documented data.

Similar roles