Senior Data Scientist, Agentic AI Systems
Axle $130K–$150K/yr
Biotechnology Research · 501-1,000 employees
About the role
Build and maintain agentic AI systems for rare disease research workflows, ensuring reliable, validated, and traceable outputs. Collaborate with NIH investigators to design data models and contribute to scientific manuscripts and conference presentations.
What they look for
Requirements
Requires a bachelor's degree in a relevant field and at least 5 years of experience in production software, including 2 years specifically with LLM-powered applications. Candidates must be able to obtain and maintain a Public Trust Security clearance.
Benefits
Full description
(ID: 2026-3432)
Axle is a bioscience and information technology company that offers advancements in translational research, biomedical informatics, and data science applications to research centers and healthcare organizations nationally and abroad. With experts in biomedical science, software engineering, and program management, we focus on developing and applying research tools and techniques to empower decision-making and accelerate research discoveries. We work with some of the top research organizations and facilities in the country including multiple institutes at the National Institutes of Health (NIH).
Benefits We Offer:
- 100% Medical, Dental & Vision Coverage for Employees
- Paid Time Off and Paid Holidays
- 401K match up to 5%
- Educational Benefits for Career Growth
- Employee Referral Bonus
- Flexible Spending Accounts:
- Healthcare (FSA)
- Parking Reimbursement Account (PRK)
- Dependent Care Assistant Program (DCAP)
- Transportation Reimbursement Account (TRN)
Axle is seeking a Senior Data Scientist, Agentic AI Systems to join our vibrant team supporting rare disease research at the National Institutes of Health (NIH). This is a Remote position within the United States.
Position Summary
Roughly 25 to 30 million people in the United States live with a rare disease. There are somewhere between 7,000 and 10,000 distinct rare conditions, and the large majority have no FDA-approved treatment.
Research on these conditions keeps running into the same obstacles. Published evidence for any one disease is thin and scattered across sources. The same clinical finding gets written down a dozen different ways depending on who recorded it. And the people with the most at stake, patients and their families, are usually the least equipped to read the specialist literature written about their own condition.
Large language models are well suited to this class of problem, and the research programs we support are investing in applying them carefully. In this role you will build the conversational AI systems that sit between a person and the research infrastructure. These are multi-turn workflows that ask sensible follow-up questions in plain language, capture the answers as validated structured data, and hand that structure off to the searches and analyses doing the scientific work. The emphasis is on systems people can rely on, which in practice means confirming every interpretation before it is saved and logging every automated decision so that it can be reviewed later.
This is a senior individual contributor position. You will own major components from design through deployment, work directly with NIH program staff, clinical geneticists, and rare disease information specialists, and help set the engineering standards for how AI gets applied on this team.
Core Responsibilities
- Build agentic AI systems for rare disease research workflows. This includes the conversation logic, the rules that decide when enough information has been gathered, and the confirmation steps that catch a misreading before it reaches anything downstream.
- Model outputs in Pydantic and use structured output and tool calling, so that every field a model produces is typed, validated, and traceable back to its source.
- Write, version, and regression test the prompts behind clinical and scientific reasoning tasks. Prompts and output schemas are treated as code here, with tests to match.
- Build evaluation for tasks that have no single right answer. Golden sets, offline regression suites, and model-based graders all have a place, and the results should be good enough to decide what ships.
- Keep multi-step LLM workflows responsive under load. This covers async design, concurrency limits, streaming partial results to the client, and timeout and failure handling that holds up in production.
- Log what the system does and why. Request identifiers, latency, errors, and the reasoning behind each automated choice all need to be captured, so that staff can review an AI-assisted result instead of taking it on faith.
- Work out what researchers, clinicians, and patient communities need, and turn it into data models and system behavior.
- Write the work up. You will contribute to manuscripts, conference abstracts, and posters with NIH investigators, and you will be credited as an author on work you helped produce.
Required Qualifications
- Bachelor’s degree in Data Science, Computer Science, Bioinformatics, Biomedical Informatics, or a related field. An advanced degree is preferred. We will consider equivalent professional experience in place of a degree.
- At least 5 years building and operating production software or data systems. At least 2 of those years should involve shipping LLM-powered applications (agents, retrieval, or evaluation) that people depend on. We weigh depth in agentic workflow engineering more heavily than total years.
- Experience with structured output and tool or function calling, meaning you have constrained a model to a typed schema and validated what came back.
- Experience evaluating systems that have no single right answer, using golden sets, offline regression suites, or model-based graders to decide whether a change was an improvement.
- Ability to own a service end to end, from schema design through deployment and operation.
- Ability to obtain and maintain a Public Trust Security clearance.
Technical Skills
- Python, with FastAPI, Pydantic, and pytest.
- LLM application engineering: provider APIs and gateways, prompt and context design, structured generation, tool use, and tracing.
- PostgreSQL, including work with embeddings or vector search alongside relational data.
- Asynchronous and concurrent Python, plus streaming results to a client.
- Containers and Kubernetes, enough to ship, debug, and operate a service on infrastructure you do not administer.
- Git-based collaboration and CI/CD in a shared codebase.
Preferred Skills
- A typed agent framework such as Pydantic AI, LangGraph, or the OpenAI or Anthropic agent SDKs, and MCP for tool integration.
- LLM tracing and evaluation tooling such as Langfuse, LangSmith, Arize Phoenix, or Braintrust.
- Serving open-weight models in production with Ollama or vLLM behind a gateway such as LiteLLM.
- Biomedical ontologies and controlled vocabularies, including MONDO, HPO, UMLS, MeSH, and other OBO Foundry resources, along with comfort working through term hierarchies, synonyms, and cross references.
- Background in rare disease, clinical genetics, or translational research.
- Experience working alongside clinicians, curators, or patient advocacy organizations, and translating their vocabulary into a data model that holds up.
- Published or presented work that explains your engineering to people who did not build it. Peer-reviewed papers, conference talks, preprints, technical blog posts, and public open source contributions all count.
- Prior or current NIH experience.
We are looking for an engineer first. If you have shipped LLM systems that people depend on and have never opened an ontology file, we want to hear from you. Rare disease and ontology background is useful but not required, and we expect to teach the domain to whoever we hire. Candidates who meet the required qualifications and none of the preferred ones are encouraged to apply.
Disclaimer: The above description is meant to illustrate the general nature of work and level of effort being performed by individuals assigned to this position or job description. This is not restricted as a complete list of all skills, responsibilities, duties, and/or assignments required. Individuals may be required to perform duties outside of their position, job description or responsibilities as needed.
The diversity of Axle’s employees is a tremendous asset. We are firmly committed to providing equal opportunity in all aspects of employment and will not tolerate any illegal discrimination or harassment based on age, race, gender, religion, national origin, disability, marital status, covered veteran status, sexual orientation, status with respect to public assistance, and other characteristics protected under state, federal, or local law and to deter those who aid, abet, or induce discrimination or coerce others to discriminate.
Accessibility: If you need an accommodation as part of the employment process please contact: careers@axleinfo.com
This role has a market-competitive salary with an anticipated base compensation range listed below. Actual salaries will vary depending on a candidate’s experience, qualifications, skills, and location.
Salary Range
$130,000—$150,000 USD
Similar roles
-
Senior Lead Data Scientist
Compass Aventura, Florida, United States · $204K–$227K/yr
-
Senior Data Scientist, Analytics & Informatics
UPMC Pittsburgh, Pennsylvania, United States
-
Sr Staff Data Scientist, New Verticals
Flex New York, New York, United States · $184K–$225K/yr
-
Data Scientist - CDI
Descartes Underwriting Paris, Ile-de-France, France
-
Senior Data Scientist
Givelify United States
-
Lead Data Scientist
McGraw Hill LLC. United States