P

ML Data Platform Engineer

Parisi Labs, Inc. New York, New York, United States · $175K–$245K/yr

3 d ago
Remote Senior (5-10 yrs) Full-time United States
Create a free account to apply — email only, no card. You can also save this posting or score it against your profile with AI.

About the role

You will build and maintain robust data ingestion, validation, and observability systems to support machine learning models and product deployments. The role involves collaborating with researchers and engineers to define data contracts and create reliable, scalable data foundations.

What they look for

Data Engineering Machine Learning Software Engineering Data Infrastructure Data Modeling Time-dependent Data Data Pipelines Data Validation Observability Data Contracts System Debugging Python SQL Data Lineage Production Systems

Requirements

Candidates must have a proven track record of designing and operating production data systems that support engineering or research workloads. Strong programming, data modeling, and debugging skills are required, along with experience managing time-dependent data and complex data pipelines.

Benefits

Equity

Full description

About Parisi Labs

Parisi Labs is building foundational world models for physical industry. We are developing models that learn how complex physical systems behave and reuse that understanding across forecasts, scenarios, and operational decisions.

Energy is our first proving ground. We combine historical and live data with operational context, bringing together machine learning research, data infrastructure, and software engineering to turn advances in modeling into useful technology for energy operators.

We are a small technical team working directly with the founders on our core models, systems, and products.

About the role

We are looking for an ML Data Platform Engineer to make the data behind our models, products, and customer deployments dependable, understandable, and easy to use.

This role sits where data engineering meets machine learning. You will turn messy, changing real-world sources into durable datasets and interfaces that researchers and engineers can trust. Your work will support both public data and customer-authorized operational data.

The goal is not to build a large platform for its own sake. It is to make each new model, product capability, and data source faster to bring online without compromising correctness. You will own the shared data foundations, working with the applied-AI engineer on model requirements and the product engineer on application needs.

What you'll own

  • Build and improve ingestion, backfills, validation, and observability for high-volume, time-dependent data.
  • Define clear data contracts and point-in-time semantics for model training, evaluation, and product use.
  • Create reusable workflows for bringing public and customer-authorized sources into the system.
  • Build quality, lineage, freshness, and access controls that make data trustworthy in repeated use.
  • Develop efficient datasets and query interfaces for machine-learning and product workloads.
  • Diagnose whether failures originate in source data, transformations, model inputs, or serving systems.
  • Work with researchers and engineers to turn recurring data requirements into reliable software rather than manual projects.
  • Decide which abstractions should become shared infrastructure and which should remain purpose-built.

What we're looking for

  • A record of designing, building, and operating production data systems that researchers or engineers depend on.
  • Strong programming, querying, and data-modeling skills, with the ability to write maintainable, tested production software.
  • An understanding of time-dependent data correctness, including backfills, revisions, freshness, and point-in-time availability.
  • Experience supporting ML training and evaluation, or similarly demanding data-intensive product workloads.
  • The ability to design practical data contracts, validation, observability, and access controls without overbuilding the platform.
  • Strong debugging and collaboration skills, including the ability to trace failures across systems and explain data limitations clearly.

First 90 days

  • 30 days: Understand the data lifecycle behind Ask The Grid and our ML work, and identify the most consequential reliability and usability gaps.
  • 60 days: Ship a reusable ingestion, backfill, validation, or dataset capability used in active product or research work.
  • 90 days: Own a dependable end-to-end data workflow, with documented contracts, quality checks, and clear operational visibility.

Why join

The quality of our models and products depends on the quality of their underlying data. You will shape that foundation early, working directly with researchers and engineers who use it, and see your work support new experiments, product capabilities, and customer deployments.

Location and compensation

Location: New York City or Boston/Cambridge. This is a hybrid role — we expect in-person collaboration 3 days per week in person.

Salary range: $175,000–$245,000 USD.

Equity: meaningful early-company equity.