About the role
You will build the data foundation for AI drug discovery by developing schemas, QC processes, and annotation pipelines for spectral datasets. You will also collaborate with machine learning and chemistry teams to validate workflows and ensure data integrity for molecular structure prediction models.
What they look for
Requirements
A PhD or equivalent industry experience in analytical mass spectrometry or cheminformatics is required. You must have hands-on experience with LC-MS/MS data, Python scripting, and structural representation methods like SMILES.
Full description
Overview
Novogaia is an applied AI drug discovery company. We build machine learning systems that decode the chemistry of natural organisms, starting with fungi, to find the next generation of medicines.
We are a small team of AI engineers, computational biologists and chemists building foundation models for molecular structure prediction from mass spectrometry data. We are seeking a computational data scientist who can define what a trustworthy spectral dataset looks like, and build the schema, QC gates, and annotation process that gets us there. This is a data and cheminformatics role, not a wet-lab role: though you'll work closely with analytical chemistry collaborators who run the instruments.
The Role
- Developing a deep understanding of Novogaia's compound library, spectral data, and how both feed into our models
- Working closely with our machine learning team to assess model training/validation leakage, de-duplicate against public datasets, and help select representative subsets for benchmarking
- Working with analytical chemist collaborators to route ambiguous or high-value spectra for expert review, so results can be compared systematically against model predictions
- Building and evaluating predictive models to infer molecular properties from molecular structure
- Validating processing workflows for raw spectra arriving from analytical partners, including validating metadata, batch tracking, versioned releases, maintaining provenance, and licensing tags at the record level
In your first year, you'll build the data foundation everything else depends on: a documented schema, a QC process, and a dataset our modeling and evaluation teams can trust.
What We Require
- Background in analytical mass spectrometry or cheminformatics (PhD or equivalent industry experience), ideally with exposure to natural products or small-molecule drug discovery
- Deep familiarity with structural representation methods such as SMILES, SMARTS, SAFE, as well molecular fingerprinting and structural embedding
- Familiarity with statistics and ML concepts, for close collaboration with the rest of the team
- Hands-on experience working with LC-MS/MS data and standard formats and open-source tools (e.g. mzML, MSConvert) and spectral databases (e.g. GNPS, MassBank, MoNA)
- Scripting ability in Python for developing algorithms and Nextflow/Snakemake for building pipelines
- Familiarity with utilizing relational database schemas and ontologies to host and structure the variety of datatypes and datasets you will encounter
What We Value
- Ability to intuitively interpret and assess mass spectrometry data and corroborate automated QC checks
- Strong scientific judgment and a willingness to flag data that isn't ready, even under deadline pressure
- Ability to turn "make this dataset AI-ready" into a concrete schema, checklist, and pipeline
- Motivation to build data infrastructure other people will confidently rely on
- Experience with natural product dereplication and compound classification
- Curiosity, low ego, and a willingness to get close to the modeling and evaluation side of the work, even if it's outside your original training
Similar roles
-
Staff Data Scientist, Sales Analytics
10x Genomics Pleasanton, California, United States · $198K–$268K/yr
-
Data Scientist
CCS INC Plano, Texas, United States · $129K–$139K/yr
-
Data Scientist Talent Network
Jobgether India · $208K–$312K/yr
-
Data Scientist-Mid
General Dynamics Information Technology Tampa, Florida, United States · $115K–$155K/yr
-
Data Scientist (Multiple openings) in Minneapolis, MN.
U.S. Bank National Association Minneapolis, Minnesota, United States · $142K–$155K/yr
-
Principal Data Scientist – Financial Crime Model Validation
Commonwealth Bank Sydney, New South Wales, Australia