Lead Data Scientist, AI Labs
NPR Washington, District of Columbia, United States · $164K–$201K/yr
Broadcast Media Production and Distribution · 1,001-5,000 employees
About the role
The Lead Data Scientist will oversee the selection, fine-tuning, and scaling of foundational AI models to enhance content metadata and audience personalization. They will also design automated machine learning pipelines and establish rigorous evaluation frameworks to ensure factual accuracy and ethical compliance.
What they look for
Requirements
Candidates must have 8+ years of professional experience in Data Science, Machine Learning, or NLP with a proven track record of shipping production-grade systems. An advanced degree in Computer Science, Data Science, or Machine Learning is required, with a preference for a PhD.
Benefits
Full description
OVERVIEW
A thriving, mission-driven multimedia organization, NPR produces award-winning news, information, and music programming in partnership with hundreds of independent public radio stations across the nation. The NPR audience values information, creativity, curiosity, and social responsibility – and our employees do too. We are innovators and leaders in diverse fields, from journalism and digital media to IT and development. Every day, our employees and member stations touch the lives of millions worldwide.
Across our organization, we’re building a workplace where collaboration is essential, diverse voices are heard, and inclusion is the key to our success. We are committed to doing the right thing in our journalism and in every role at NPR. This means that integrity, adherence to our ethical standards, and compliance with legal obligations are fundamental responsibilities for every employee at NPR.
Intro to Position
As the Lead Data Scientist for the AI Labs team, you will serve as the technical and ethical anchor for NPR's artificial intelligence initiatives. You will lead data science expertise for a content metadata overhaul to power new audience-focused personalization engines. Rather than building models from scratch, you will focus on fine-tuning, evaluating, and scaling existing foundational models, as well as deploying machine learning algorithms for public media.. You will collaborate extensively across the organization to align architectures, ensure secure cloud deployments, and protect intellectual property. This position demands a high focus on accuracy, journalistic ethics, data privacy, and responsible scaling.
Responsibilities
- Lead the selection, fine-tuning, and optimization of open-source and proprietary LLMs tailored to NPR’s unique content voice.
- Design and architect automated machine learning pipelines to transform decades of unstructured audio, transcripts, and text to support automated semantic metadata generation.
- Collaborate with the product, design and engineering teammates to translate user needs into production-ready data science workflows.
- Architect recommendation frameworks that leverage enriched metadata to drive deep, style-based audience personalization while preserving editorial curation.
- Partner with Data Products to construct clean, self-service data pipelines and audience analytics models inside BigQuery.
- Support newsroom research through the prototyping, validation, and development of API-driven tooling and secure database search models.
- Establish strict evaluation, testing, and benchmarking frameworks to guarantee model outputs meet NPR’s standards for factual accuracy and neutrality.
- Proactively identify, audit, and mitigate algorithmic bias in metadata generation and audience discovery systems.
- Ensure all AI applications scale securely and cost-effectively, balancing computational efficiency with rigorous data privacy guardrails.
- Collaborate with growth and data platforms to leverage content metadata for user lifecycle retention and smart audience segmentation.
The above duties and responsibilities are not an exhaustive list of required responsibilities, duties and skills. Other duties may be assigned, and this job description is subject to change at any time.
Minimum Qualifications
- 8+ years of professional experience in Data Science, Machine Learning, or Natural Language Processing (NLP) shipping production-grade systems.
- Proven track record of applying, fine-tuning, and evaluating Large Language Models (LLMs) and foundational architectures.
- Experience with Automatic Speech Recognition (ASR), diarization, and audio preprocessing pipelines.
- Demonstrated experience designing and maintaining large-scale data architectures, vector databases, and semantic search pipelines.
- Hands-on experience designing, deploying, and optimizing production-grade recommendation engines or personalization systems at scale.
- Practical, hands-on experience building machine learning workflows within major cloud environments (Azure, AWS or GCP).
- Experience successfully navigating matrixed, cross-disciplinary collaboration between technical engineering teams and non-technical stakeholders.
- Ability to identify common AI pitfalls, such as hallucinations or formatting errors, and design robust algorithmic workarounds.
Preferred Qualifications
- Experience working with large volumes of unstructured multimedia, digital audio processing, or automated speech-to-text workflows.
- Prior experience working within a media organization, digital newsroom, or public service institution.
- Experience integrating automated LLM evaluation gates and regression testing directly into CI/CD deployment pipelines
Required Skills/Competencies
- Deep technical expertise in Python, ML and agentic engineering frameworks, vector search engines and two-stage retrieval architectures.
- Strong proficiency in SQL and cloud data warehouse ecosystems (such as BigQuery or Snowflake).
- Deep understanding of Retrieval-Augmented Generation (RAG) patterns and semantic search tools.
- Professional-level familiarity with model evaluation methodologies, prompt engineering techniques, and API integration workflows.
- Strong commitment to algorithmic ethics, data privacy compliance, and methods for identifying or mitigating bias.
Education Requirements
- Advanced degree (PhD preferred) in Computer Science, Data Science or Machine Learning
- Advanced academic research in relevant fields
Work Location & Requirements
- Hybrid Permitted:
- This is a hybrid permitted role. Some aspects of this role require duties better performed at an NPR facility. The employee will be required to be in the office at the [Washington, D.C. or New York City] location at least two to three days per week.
Job Type
- This is a full-time, exempt position.
Compensation
Salary Range: The U.S.-based anticipated salary range for this opportunity is $164,000 – $201,000 plus benefits. The range displayed reflects the minimum and maximum salaries NPR expects to provide for new hires for the position across all US locations.
NPR Benefits: NPR provides comprehensive benefits for employees and dependents. Regular, full-time employees scheduled to work 30 hours or more per week are eligible to enroll in NPR’s benefits options. Benefits include access to health and wellness, paid time off, and financial well-being. Plan options include medical, dental, vision, life/ accidental death and dismemberment, long-term disability, short-term disability, and voluntary retirement savings to all eligible NPR employees.
Does this sound like you? If so, we want to hear from you.
#LI-HYBRID
The range displayed reflects the minimum and maximum salaries NPR expects to provide for new hires for the position across all US locations.
NPR Pay Range
$164,000—$201,000 USD
NPR is an Equal Opportunity Employer. NPR is committed to being an inclusive workplace that welcomes diverse and unique perspectives, all working toward the same goal – to create a more informed public. Qualified applicants receive consideration for employment without regard to race, color, ethnicity, national origin, ancestry, age, religion, religious belief, sex (including pregnancy, childbirth and related medical conditions, lactation, and reproductive health decisions), sexual orientation, gender, gender identity or expression, transgender status, gender non-conforming status, intersex status, sexual stereotypes, nationality, citizenship status, personal appearance, marital status, family status, family responsibilities, military status, veteran status, mental and physical disability, medical condition, genetic information, genetic characteristics of yourself or a family member, political views and affiliation, unemployment status, protective order status, status as a victim of domestic violence, sexual assault, or stalking, or any other basis prohibited under applicable law.
If you are a person with a disability needing assistance with the application process, please reach out to employeerelations@npr.org.
You may read NPR’s privacy policy to learn about how NPR may handle information you submit with any application.
Want more NPR? Explore the stories behind the stories on our NPR Extra blog. Get social with NPR Extra on Facebook and Instagram. Find more career opportunities at NPR.org/careers.
Similar roles
-
Sr. Data Scientist
M J Brunner Inc Pittsburgh, Pennsylvania, United States
-
Senior Data Scientist, Membership
Chime Financial, Inc San Francisco, California, United States
-
Senior Data Scientist, Growth Product
Chime Financial, Inc San Francisco, California, United States · $133K–$185K/yr
-
Staff Data Scientist
Robots and Pencils Canada
-
Staff Data Scientist
Tenpo Monte Patria, Coquimbo Region, Chile
-
Senior Associate, Data Scientist - PXT Analytics
JPMorgan Chase & Co. Jersey City, New Jersey, United States · $124K–$170K/yr