The future is the future - thought about access to data in publications
Donat Agosti · Planning the Libroscope: creating research ready biodiversity from scientific publications
The short versionLiberating data from literature works, but the future lies in use-case-driven, digital-first publishing that also targets the undescribed majority of species.
What this was about
Donat Agosti introduced the Libroscope and the Disentis roadmap goal of making 100% of major biodiversity publications open, FAIR, machine-actionable and AI-ready by 2035, using a use-case-driven approach. He described Plazi's current pipeline, in which PDFs or XML are processed so that data reach GBIF, Zenodo and dashboards within a day of publication, with provenance down to the exact position in the text, and noted 1.2 million treatments extracted and a steady ~10,000 new species a year across roughly 200 journals processed. He argued the community should also move beyond extracting from PDFs toward digital-first description of unknown and dark taxa.
Why it matters. Literature holds much of what we know about biodiversity yet remains the least open data type; this talk frames a roadmap for making it trusted, AI-ready data.
In the room
- Estimates: 500 million printed pages, ~2,000 taxonomic journals, 20-50 million treatments, and only 10-30% of species data accessible.
- The 2024 Disentis meeting produced a roadmap to make all major biodiversity publications open and FAIR by 2035.
- The Libroscope is use-case-driven: target specific data rather than everything at once.
- Around 30 GBIF hosted portals are dashboards from publications; material citations link back to exact text positions via persistent identifiers.
- Zenodo communities (e.g. for fruit flies) and the Biodiversity Literature Repository collect publications for processing.
- Lag from publication to GBIF has fallen from years to under a day in the best case.
- 1.2 million treatments extracted; ~200 journals processed over 10 years show a flat rate of ~10,000 new species a year.
Notable moments
Transcript
Automatically generated captions can contain mistakes, especially in names and technical terms. Times are relative to the room recording.


