S

Data Engineer

Stealth - AI Agents for Healthcare San Francisco, California, United States

Yesterday
data-engineer Senior (5-10 yrs) Full-time United States
Create a free account to apply — email only, no card. You can also save this posting or score it against your profile with AI.

About the role

You will build and own the foundational data systems to ingest, model, and serve reliable healthcare data across the company. Additionally, you will leverage AI to automate data acquisition and reconciliation while setting the roadmap for data infrastructure.

What they look for

Data Engineering Python SQL Data Modeling Data Warehousing Databricks Airflow Dagster Dbt Data Ingestion Data Pipeline Design AI Integration System Architecture Data Quality Data Observability Cloud Infrastructure

Requirements

The ideal candidate is an end-to-end data engineer with strong experience in Python, SQL, and production data systems. You must be comfortable working with messy, complex healthcare data and capable of designing scalable, durable architectures from scratch.

Full description

About Us

We are reimagining how people buy and navigate healthcare.

More than 80% of Americans receive health insurance through an employer or purchase coverage through individual and Medicare exchanges. Yet the experience remains overpriced, opaque, and inefficient. We replace the traditional brokerage with an AI-native platform built to streamline this process while helping people choose plans, find care, schedule appointments, call insurance, and fight bills.

Our goal is not to build a better brokerage. It is to redefine how healthcare is bought and managed across America.

Why Join Us

  • Scaling quickly: We are growing quickly and looking for zero-to-one builders who want an outsized opportunity to learn and shape the company’s future.
  • High impact: Healthcare costs are rising at record rates, while the user journey has never been worse. Join us as we rebuild an industry.
  • Best-in-class talent: Work alongside teammates from leading tech companies including Uber, Faire, Airbnb, Rippling, Google, Forward, and Collective Health, as well as brokerage industry leaders including Newfront, WTW, Marsh, and Sequoia. Our team works together in person in San Francisco.
  • Well funded: Seed-stage startup incubated alongside the healthcare and AI roll-up teams at General Catalyst.

About the Role

Our product depends on bringing together data that has never previously lived in one system: insurance plans and networks, provider and pricing data, employer and member information, claims and medical records, and the activity generated across our own products and operations.

We are looking for our first Data Engineer to build the foundation that makes this information reliable and usable across the company. This person will own the systems that ingest data from APIs, files, vendors, customer systems, public sources, and legacy portals; model it into durable representations; and make it available to our products, AI agents, operators, and business systems.

This is a hands-on role with broad ownership across the data stack. You will make early architectural decisions, build the first versions yourself, and determine where rigor matters immediately and where simple systems can evolve. You will also help define how we use AI in data engineering to acquire, interpret, reconcile, and operate data more effectively.

The right candidate has strong data engineering judgment, enjoys building from zero, and is excited to work with messy healthcare data where accuracy matters.

What You’ll Do

  • Build our data foundation: Own the systems that move, transform, organize, and serve data across the company, including the data warehouse (e.g. Databricks), ingestion and export infrastructure, orchestration and scheduling (e.g. Airflow or Dagster), transformation and modeling tools (e.g. dbt), semantic layers, data catalogs, quality, and observability.
  • Unify fragmented healthcare data: Ingest and connect data from paid vendors, carriers, providers, government sources such as CMS, customer and benefits administration systems, public files, APIs, legacy portals, and our own applications. Build durable workflows for large files, changing schemas, incomplete records, and sources that were never designed to work together.
  • Make the data trustworthy: Design canonical models for plans, networks, providers, employers, members, claims, and other core concepts. Preserve historical state, lineage, provenance, and source-specific nuance so our products, agents, operators, and customers can rely on the results.
  • Put data into production: Make curated data available wherever it creates value, including our backend, member and employer experiences, AI agents, brokerage workflows, CRM, reporting, and other operational tools. Build reliable systems for moving data both into and out of the warehouse.
  • Use AI to improve data engineering: Apply models and agents to acquire information from portals and older websites, interpret unfamiliar files, map changing schemas, reconcile conflicting sources, investigate quality issues, and automate work that has traditionally been manual.
  • Set the roadmap and standards: Work across product, engineering, brokerage, and operations to identify the highest-leverage data capabilities, prioritize the roadmap, and decide where we need durable infrastructure now and where a simpler solution can evolve.

Who You Are

  • AI-native: You think from first principles about how models and agents can change data engineering, from acquiring data in legacy systems to interpreting files, mapping schemas, and investigating failures. You know when to use AI, when deterministic systems are safer, and how to build the evaluations and review loops needed for trustworthy results.
  • End-to-end data engineer: You have built and operated production data systems across ingestion, warehousing, orchestration, modeling, quality, and serving. You are hands-on in SQL and Python and comfortable choosing and owning the surrounding tools.
  • Comfortable starting from zero: You can turn broad goals into a clear sequence of work and take responsibility from initial design through reliable production operation.
  • Comfortable with messy data: You have worked with data from external organizations, not only clean application events. You understand temporal modeling, idempotency, schema evolution, lineage, provenance, and how to establish trust when sources are incomplete or contradictory.
  • Practical architect: You can make important technical decisions without overbuilding. You know which choices will be difficult to reverse, where correctness is essential from the beginning, and where a simple first version is appropriate.
  • Operationally minded: You think beyond dashboards and analytical queries. You are comfortable designing data systems that directly support applications, automated workflows, customer-facing experiences, and business operations.
  • Able to work across disciplines: You enjoy working with engineers, operators, and industry experts to turn complicated real-world concepts into clear models, interfaces, and systems.

Similar roles