Pedestal Health

Principal Data Scientist

Pedestal Health Research Triangle Park, North Carolina, United States

Biotechnology Research · 51-200 employees

4 h ago
data-scientist Senior (5-10 yrs) Full-time United States
Create a free account to apply — email only, no card. You can also save this posting or score it against your profile with AI.

About the role

You will design and scale AI-powered internal tools to improve analytic workflows and data quality across the organization. This role involves technical leadership, mentorship, and cross-functional collaboration to ensure data trustworthiness and operational efficiency.

What they look for

Data science Python SQL Large language models Data quality Machine learning Cloud data warehouse Snowflake Technical leadership Mentorship Data infrastructure Healthcare data EHR Claims data Prompt engineering System architecture

Requirements

Candidates must hold an MS or PhD in a quantitative field and possess significant experience with real-world healthcare data. You should have a proven track record of building internal tools and platforms, along with strong proficiency in Python, SQL, and LLM implementation.

Benefits

Health insurance Dental insurance Vision insurance 401(k) with company match Generous PTO Company holidays Paid parental leave

Full description

The Role

We are seeking a Principal Data Scientist to join Pedestal Health's Quantitative Sciences (QS) organization. In this role, you will build the infrastructure and methods that make AI-assisted analytics and real-world data quality work fast, scalable, and genuinely trustworthy.

Your work will span two connected areas. The first is AI-enabled analytics: building and maintaining AI tooling and workflows that let our team identify cohorts, review analysis code, and carry out recurring analytic work faster and more consistently. The second is AI-enabled data quality: designing automated and agentic approaches that scale across schemas, sites, and data refreshes, and that surface issues before they reach an analysis or a client.

This is a role about building capability, not about producing analyses. You will design, build, and scale the internal tools that change how our teams work with real-world data — and you will own them as products, with users, versions, quality standards, and a roadmap. This is a hands-on role that combines individual technical contribution with technical leadership, including mentorship and review of other data scientists' work. You will partner closely with Product, Engineering, Medical, and Commercial teams, and report to the Head of Quantitative Sciences.

What You'll Do

Build and Scale Internal Tooling

You will make AI a dependable part of how our analytic work gets done, not an occasional shortcut.

  • Build, maintain, and improve AI-powered tooling that lets the team generate commercial cohort counts and conduct feasibility reliably and repeatably
  • Extend the same approach to other recurring analytic work, including generating and reviewing analysis code, and supporting protocol and analysis plan development
  • Gather requirements from the internal teams who depend on these tools, treat them as users, and iterate on real feedback rather than assumed needs
  • Own what keeps this tooling trustworthy over time (how it is tested against known-correct results, how updates are validated before release, and how performance is monitored as the underlying data evolves), and where human review remains mandatory
  • Scale adoption across the team through documentation, training, onboarding, and hands-on enablement
  • Help define the guardrails for AI use in client-facing work: what data may be used, how outputs are reviewed and by whom, and how provenance is recorded

AI-Enabled Data Quality

You will design the infrastructure that tells us whether our data is trustworthy, before anyone else has to find out.

  • Rethink how our data quality checks are built and run, so that assessing a new source or a refreshed schema no longer means redoing the work each time
  • Move quality assessment beyond manual review and spreadsheet outputs, toward an automated approach with a durable record of what was checked, what was found, and how it was resolved
  • Determine where AI and agentic approaches genuinely add leverage in this work, and where deterministic, reproducible checking should remain the foundation
  • Define how we will know the system is working, including whether the people who receive quality signals continue to trust them and act on them
  • Partner with Engineering on pipeline integration, orchestration, and monitoring

Cross-Functional Collaboration and Team Development

You will connect data science to the teams that build, sell, and deliver on top of it.

  • Partner closely with Product, Engineering, Medical Science, Clinical Operations, and Commercial teams
  • Advocate effectively for the quality and validation work that AI-assisted products require, including in roadmap and prioritization discussions
  • Keep pace with developments in AI tooling and methods, and actively push what proves useful into our standards, training, and tooling
  • Provide technical oversight, code review, and mentorship to other data scientists, and own the standards their work is measured against
  • Help grow the data science group as it expands

What You'll Bring

  • MS or PhD in data science, statistics, computer science, computational biology, or a related quantitative field; equivalent practical experience will be considered
  • Experience working with real-world healthcare data — EHR, claims, or registry — including direct familiarity with its structural and quality challenges
  • A track record of building internal tools, platforms, or analytic products used by other people, and owning them through multiple versions — not solely a record of delivering analyses
  • Strong programming ability in Python and SQL, with experience in a cloud data warehouse environment such as Snowflake
  • Demonstrated ownership of a data quality, testing, or observability system end to end, not only authoring individual checks
  • Practical experience building with large language models beyond prototyping, including prompt and workflow design and a clear point of view on how to evaluate whether an LLM-based system is actually working
  • Comfort working in raw, messy, poorly documented source data, and the curiosity and persistence to figure out what it actually contains
  • Experience working cross-functionally with product and engineering partners, and the ability to hold a technical line constructively when priorities compete
  • Clear written and verbal communication, including the ability to explain a method's limitations as readily as its results
  • Interest in growing into technical leadership, including mentoring and reviewing the work of other data scientists

What we offer you

  • Hybrid work — 3 days/week in our brand-new office!
  • Comprehensive health, dental, and vision for you and your family
  • 401(k) with company match
  • Generous PTO and company holidays
  • Paid parental leave

Hybrid role: Located in Research Triangle Park, North Carolina

If you are ready to be part of a team where your work truly matters—where your expertise is valued, your growth is supported, and your contributions help shape the future of healthcare—Pedestal Health is the place for you. We’re building something meaningful together, and we’d love for you to be a part of it.

Pedestal Health is an equal opportunity employer and seeks candidates from diverse backgrounds and abilities.

Similar roles