talk · Thursday 24 September · SAL C

How Literature Services can Support & Benefit from Biodiversity Publication and Data Standards?

Patrick Ruch · Digital Tools for Data Discovery, Resolution and Exchange

Recording time 8:18:44–8:29:47Open on Vimeo ↗

The short versionStandards-based concept annotation of biodiversity literature enables focused, evidence-grounded question answering that is more precise than general chatbots.

Overview

What this was about

Patrick Ruch presented BiodiversityPMC, a SIB Literature Services front end that indexes MEDLINE, an extended PubMed Central including Pensoft and other non-PubMed journals, a supplementary-data index with about 24 million OCR'd images, Plazi treatments and Zenodo. It is a life-science-focused superset of PubMed. All content is pre-annotated with about 50 terminologies (close to nine million concepts), so vernacular names, synonyms and taxa normalise to identifiers. He showed concept-normalised question answering: an extractive mode that returns facts from papers (F1 above 70% for the top answer on a SQuAD-style benchmark) and a generative RAG mode that answers from a few retrieved documents, contrasted with verbose general chatbots. The generative mode was due to become public in October.

Why it matters. Semantic normalisation of literature with shared vocabularies lets researchers and AI agents retrieve and cite specific facts across millions of papers, supplements and treatments.

Key ideas

In the room

  • BiodiversityPMC combines MEDLINE, PMC extended with journals not in PubMed (e.g. Pensoft), a supplementary data index (spreadsheets, slides, CSV and ~24 million OCR'd images), Plazi treatments and Zenodo.
  • Zenodo/grey literature content is expected to exceed one billion items by year end.
  • Coverage is a superset of PubMed, adding agriculture, ecology and environmental sciences, but narrower than OpenAlex or Google Scholar; OpenAlex does not index image content.
  • All papers are pre-annotated with ~50 terminologies (~9 million concepts), e.g. the vernacular 'raccoon dog' is normalised to Nyctereutes procyonoides in Open Tree of Life.
  • QA pipeline: normalise both collection and query to concepts, retrieve documents, then apply extractive or generative answering.
  • Extractive QA returns factoid answers from papers, with F1 above 70% for the top answer on a SQuAD-based benchmark.
  • Generative QA uses retrieval-augmented generation with a fine-tuned, low-compute model; for Collembola trophic guilds it answered 'primarily detritivores' from two documents, versus a chatbot listing every possibility.
Jump in

Notable moments

In their words

Transcript

Automatically generated captions can contain mistakes, especially in names and technical terms. Times are relative to the room recording.

Read the transcript ↓
Loading transcript…