Senior Data Scientist I
Elsevier Amsterdam, North Holland, Netherlands · €54K–€90K/yr
Information Services · 5,001-10,000 employees
About the role
You will design, build, and scale advanced AI and machine learning systems to power scientific discovery and research intelligence. This involves leading end-to-end data science initiatives, from problem framing and model development to production deployment and continuous monitoring.
What they look for
Requirements
Candidates must have significant applied experience in data science, machine learning, or NLP, along with a Master's or PhD in a quantitative field. Proficiency in Python and experience with LLMs, retrieval systems, and production-ready code development are essential.
Benefits
Full description
Senior Data Scientist
AI for Science, Research Intelligence & Knowledge Discovery
Technology – Data Science Organization
Do you want to build advanced AI that helps researchers discover, understand, and advance science?
At Elsevier, data science isn’t about building models for their own sake. It’s about doing science for science: creating intelligent systems that help the global research and healthcare communities make better decisions, uncover trusted insights, and accelerate scientific progress.
We are growing our Data Science organization within Technology, and we are looking for Senior Data Scientists who want to work at the frontier of applied AI — machine learning, natural language processing, large language models, retrieval, reasoning workflows, knowledge discovery, and state-of-the-art generative AI built specifically for science.
You will help solve some of the hardest AI problems in the world: building systems that understand scientific language, reason across trusted content, connect ideas across disciplines, and support researchers as they explore the frontiers of knowledge.
You will work with vast, complex, and heterogeneous data — from scientific publications and raw research datasets to biomedical content, citations, metadata, ontologies, taxonomies, knowledge graphs, multilingual content, and domain-specific data across the full spectrum of scientific disciplines.
This is AI with purpose: helping researchers advance science, improve health outcomes, and contribute to human progress.
About the Role
As a Senior Data Scientist at Elsevier, you will design, build, evaluate, and scale the advanced AI and machine learning capabilities that power scientific discovery, research intelligence, editorial workflows, knowledge enrichment, and decision support.
You will work across the full lifecycle of data science solutions — from problem framing, data exploration, experimentation, and prototyping through model development, evaluation, production deployment, monitoring, and continuous improvement.
You will apply a broad range of techniques, from classical machine learning and deep learning to large language models, retrieval-augmented generation, semantic search, agentic workflows, citation-aware reasoning, and evidence-grounded AI.
This role rewards both technical depth and practical judgment. You will decide when to reach for traditional machine learning, when LLM-based approaches fit best, when retrieval or knowledge systems are required, and how to combine them into reliable, scalable, and trusted product capabilities.
You will collaborate closely with engineering, product, UX, analytics, research, editorial, and subject-matter experts to turn complex scientific and business challenges into high-impact AI solutions.
What You’ll Do
You will lead and contribute to high-impact AI and data science initiatives that help people explore, understand, connect, and act on complex scientific information.
In this role you will:
- Design, build, and evaluate advanced AI, machine learning, NLP, and generative AI systems for scientific and knowledge-discovery applications.
- Develop LLM-powered research workflows, including scientific question answering, literature summarization, semantic exploration, research insight generation, citation-aware reasoning, and evidence-grounded generation.
- Build and optimize retrieval-augmented generation systems that connect large language models with trusted scientific, biomedical, technical, and scholarly content.
- Design search and retrieval pipelines using lexical, vector, semantic, hybrid, and re-ranking approaches.
- Develop intelligent capabilities for classification, entity extraction, enrichment, ranking, recommendation, summarization, machine translation, content ingestion, contextual retrieval, and decision support.
- Create agentic, multi-step AI workflows using orchestration frameworks such as LangChain, LangGraph, Haystack, or similar tools.
- Experiment with embeddings, chunking strategies, context management, prompt engineering, grounding strategies, hallucination mitigation, re-ranking, and model orchestration.
- Integrate scientific metadata, ontologies, taxonomies, knowledge graphs, citation networks, and domain-specific knowledge assets into AI workflows.
- Build robust evaluation frameworks for AI, search, retrieval, and GenAI systems, including relevance, grounding, faithfulness, hallucination detection, quality, trust, reliability, and user impact.
- Design offline evaluation methodologies, benchmark datasets, annotation strategies, and online experimentation, including A/B testing where appropriate.
- Write clean, tested, production-ready Python code and develop reusable data science packages, libraries, and pipeline components.
- Partner with engineering teams to deploy, productionize, monitor, optimize, and maintain data science systems at scale.
- Establish reporting for model and pipeline performance, including monitoring for drift, quality degradation, latency, cost, reliability, and retraining needs.
- Lead technical discovery, shape solution design, communicate trade-offs, and influence roadmap decisions.
- Mentor data scientists, share knowledge, and foster a culture of experimentation, quality, responsible AI, and continuous learning.
What Makes This Opportunity Unique
At Elsevier, you won’t be building generic AI features. You will be building AI systems for the global knowledge ecosystem.
The problems are intellectually rich, technically demanding, and deeply meaningful. Scientific information is complex, nuanced, domain-specific, multilingual, interconnected, and constantly evolving. Researchers need tools they can trust — tools that surface evidence, preserve context, cite sources, explain their outputs, and help them move faster without compromising quality.
You may work with:
- Scientific publications, abstracts, full-text content, and research metadata.
- Raw scientific datasets and structured research-data repositories.
- Biomedical, chemical, clinical, life sciences, engineering, physical sciences, social sciences, and multidisciplinary content.
- Citations, author networks, affiliations, journals, conferences, topics, concepts, and institutional data.
- Knowledge graphs, ontologies, taxonomies, vocabularies, entity networks, and semantic enrichment systems.
- Multilingual and domain-specific content requiring sophisticated ingestion, transformation, normalization, and enrichment.
- Large-scale behavioral and usage signals that sharpen discovery, relevance, personalization, and decision support.
- Generative AI systems that must be accurate, scalable, source-grounded, explainable, measurable, and trusted.
Your work will help researchers find relevant knowledge faster, discover hidden connections, generate insights, assess evidence, improve research quality, and accelerate the advancement of science.
What We’re Looking For
We are looking for senior data scientists who pair deep technical capability with curiosity, product thinking, scientific rigor, and strong collaboration. This is a role for experienced practitioners who operate with a high degree of autonomy, technical ownership, and cross-functional influence — people who can lead complex work from idea to impact, shape technical approaches, make architectural and methodological trade-offs, design rigorous evaluations, communicate clearly with stakeholders, and mentor others.
You might lead major components of a product capability, own end-to-end delivery of data science systems, drive experimentation and evaluation strategy, or set technical direction across initiatives. We will shape the role around the strengths you bring, balancing innovation with reliability, speed with quality, and model performance with trust, explainability, scalability, and real-world usefulness.
Core Experience
- Significant applied experience in data science, machine learning, artificial intelligence, NLP, information retrieval, statistics, applied mathematics, computer science, or a related quantitative field.
- A Master’s, PhD, or equivalent practical experience in Computer Science, Data Science, Artificial Intelligence, Machine Learning, NLP, Information Retrieval, Mathematics, Statistics, or a related discipline.
- Strong hands-on experience building AI, ML, NLP, GenAI, or retrieval-based systems in applied or product-oriented environments.
- Advanced Python skills, with experience writing production-quality, maintainable, tested, and well-documented code.
- A strong command of machine learning fundamentals, including classification, regression, clustering, ranking, deep learning, model evaluation, feature engineering, validation, and performance measurement.
- Experience with large language models in applied settings, including integration into workflows, evaluation of outputs, prompt engineering, grounding strategies, and responsible use.
- Experience designing or contributing to retrieval-augmented generation systems, semantic search, embeddings, vector search, hybrid retrieval, or ranking systems.
- Experience working with large-scale structured, semi-structured, or unstructured datasets, especially text-rich or content-heavy datasets.
- A strong grasp of experimentation design, evaluation frameworks, statistical analysis, quality metrics, and measurable user impact.
- Experience with modern ML and AI frameworks such as Scikit-learn, PyTorch, TensorFlow, Hugging Face, LangChain, LangGraph, or Haystack.
- Experience with data manipulation, analysis, and visualization tools such as Pandas, NumPy, SciPy, Matplotlib, Tableau, or Power BI.
- Familiarity with cloud platforms, distributed data processing, and modern development practices, including Git, CI/CD, DevOps, and collaborative software engineering.
- Strong communication and presentation skills, with the ability to explain complex technical concepts, model behavior, risks, and trade-offs to technical and non-technical stakeholders.
- A demonstrated ability to translate ambiguous requirements into practical, scalable, measurable, and high-quality solutions.
- A track record of mentoring, technical leadership, knowledge sharing, or guiding others through complex data science work.
AI, Retrieval and Modern Machine Learning Expertise
Relevant experience may include any of the following:
- Large language models and generative AI systems.
- Retrieval-augmented generation and source-grounded AI.
- Semantic search, vector search, hybrid retrieval, and ranking.
- Scientific question answering, summarization, literature exploration, and research insight generation.
- Citation-aware reasoning, evidence-based AI, and trustworthy answer generation.
- Agentic AI, multi-step workflows, tool use, orchestration, and AI assistant experiences.
- NLP techniques such as classification, entity extraction, topic modeling, named entity recognition, machine translation, summarization, clustering, and text generation.
- Deep learning, transformer models, neural networks, transfer learning, reinforcement learning, and advanced model architectures.
- Evaluation of AI-generated outputs, including grounding, faithfulness, hallucination detection, relevance, reliability, and user value.
- Knowledge graphs, ontologies, taxonomies, metadata enrichment, semantic enrichment, and entity resolution.
- Responsible AI practices, including robustness, transparency, quality, reproducibility, and trust.
Tools and Technologies
Experience with some of the following is valuable:
- Python, SQL, Git, CI/CD, DevOps, unit testing, and object-oriented development.
- Scikit-learn, PyTorch, TensorFlow, Hugging Face, LangChain, LangGraph, Haystack, or similar frameworks.
- Databricks, Spark, Hadoop, OpenSearch, vector databases, or other large-scale data and retrieval platforms.
- AWS, Azure, Bedrock, SageMaker, or other cloud-based AI and ML services.
- MLflow, Kubeflow, SageMaker, or other MLOps and model-lifecycle tools.
- REST APIs, microservices, relational databases, document databases, JSON, XML, and semi-structured data formats.
- Experiment tracking, model registries, reproducible workflows, benchmark suites, annotation pipelines, and evaluation dashboards.
- Performance optimization techniques such as parallelization, multi-threading, batching, caching, latency optimization, and cost-aware model deployment.
- Agile delivery practices and collaborative development environments such as Jira, GitHub, or GitLab.
Nice to Have
- Experience with scientific, scholarly, biomedical, clinical, chemical, academic, publishing, or other knowledge-intensive data.
- Experience building AI assistants, conversational AI systems, research copilots, or intelligent workflow tools.
- Experience with large-scale search, ranking, recommendation, or knowledge-discovery systems.
- Experience with citation-aware, source-grounded, evidence-based, or high-trust AI systems.
- Experience with knowledge graphs, ontologies, taxonomies, vocabularies, metadata enrichment, or semantic search.
- Experience building or evaluating agentic RAG systems.
- Experience with multilingual data ingestion, preprocessing, transformation, enrichment, or machine translation.
- Experience productionizing AI or ML systems, including deployment, monitoring, drift detection, automated retraining, performance reporting, and maintenance.
- Experience optimizing systems for scale, latency, reliability, cost, or throughput.
- Experience in regulated, high-trust, or content-rich domains.
- Publications, patents, open-source contributions, or applied research in NLP, information retrieval, search, machine learning, or generative AI.
- Interest in emerging paradigms such as AI-assisted development, spec-driven development, human-in-the-loop evaluation, and responsible AI at scale.
Why Join Us
Because this is where advanced AI meets one of the most important missions in the world: advancing science.
At Elsevier, your work helps researchers save time, uncover evidence, identify patterns, connect concepts, and make better decisions. You will help build systems that make scientific knowledge more discoverable, trustworthy, and actionable.
You will have the opportunity to:
- Tackle some of the hardest AI problems in science and knowledge discovery.
- Build advanced AI systems using vast, rich, heterogeneous, and intellectually challenging scientific data.
- Combine deep machine learning expertise with state-of-the-art LLMs, retrieval, RAG, knowledge graphs, and reasoning workflows.
- Create trusted AI capabilities that support researchers, authors, editors, clinicians, institutions, and decision-makers around the world.
- Shape how AI is designed, evaluated, deployed, and governed in high-impact knowledge environments.
- Collaborate with talented colleagues across data science, engineering, product, UX, analytics, research, editorial, and domain expertise.
- Mentor others, influence technical direction, and contribute to a culture of quality, innovation, and responsible AI.
- Help researchers advance science and contribute to human progress.
Work in a Way That Works for You
We promote a healthy work/life balance and support flexible working. We know people do their best work when they have the flexibility, trust, and support they need to thrive.
We offer initiatives and benefits that may include wellbeing support, flexible working arrangements, shared parental leave, study assistance, sabbaticals, and ongoing opportunities for learning and development.
Working With Us
We are an equal-opportunity employer committed to helping you succeed.
You will find an inclusive, collaborative, agile, innovative, and supportive environment where everyone has a part to play. We value diverse perspectives, thoughtful debate, practical problem-solving, and colleagues who care deeply about what they do and how they do it.
About Elsevier
Elsevier is a global leader in information and analytics. We help researchers and healthcare professionals advance science and improve health outcomes for the benefit of society.
Building on our publishing heritage, we combine quality information, vast datasets, advanced analytics, and innovative technologies to support visionary science and research, health education, interactive learning, and exceptional healthcare and clinical practice.
At Elsevier, your work contributes to the world’s grand challenges and a more sustainable future. We harness technology to support science and healthcare in partnership with the communities we serve.
Together, we create possibilities. Join us.
If performed in NLD Amsterdam (Radarweg), the base pay range is €53,800 - €89,900.
This job may be subject to a collective labor agreement in the Netherlands. Please consult with the hiring team for further details.
We know your well-being and happiness are key to a long and successful career. We are delighted to offer country specific benefits. Click here to access benefits specific to your location.
We are committed to providing a fair and accessible hiring process. If you have a disability or other need that requires accommodation or adjustment, please let us know by completing our Applicant Request Support Form or please contact 1-855-833-5120.
Criminals may pose as recruiters asking for money or personal information. We never request money or banking details from job applicants. Learn more about spotting and avoiding scams here.
Please read our Candidate Privacy Policy.
We are an equal opportunity employer: qualified applicants are considered for and treated during employment without regard to race, color, creed, religion, sex, national origin, citizenship status, disability status, protected veteran status, age, marital status, sexual orientation, gender identity, genetic information, or any other characteristic protected by law.
USA Job Seekers:
EEO Know Your Rights.
Similar roles
-
Data Scientist III
WARRANT TECHNOLOGIES Crane, Indiana, United States
-
Data Scientist, AI/ML Model Quality
Apple San Diego, California, United States
-
Sr. Data Scientist
ExtraHop United States · $165K–$180K/yr
-
Principal Data Scientist
Pedestal Health Research Triangle Park, North Carolina, United States
-
Staff Data Scientist – AI/ML, Generative AI & Agentic Systems
Sandisk Milpitas, California, United States · $117K–$198K/yr
-
Data Scientist
Sierra Management And Technologies Inc California, Maryland, United States · $150K–$170K/yr