Explore the conversation.

Building a New Foundation for the Biodiversity Heritage Library
BHL has survived losing its Smithsonian home by becoming an independent, globally governed consortium, but its future is still not secure and it needs community support.
Nicole Kearney ↗
A closer look at steps needed to extract research-ready biodiversity data from legacy literature
Legacy literature data can only be unlocked by combining better OCR/NLP with taxonomic rigour that accounts for synonymy, shifting concepts, coreference and mixed languages.
Gerwin Kasperek ↗
BHL and the Planetary Knowledge Base: Structuring Biodiversity Literature for the Next Generation
NHM is re-OCRing all of BHL with modern models and extracting entities into a knowledge graph, with confidence gating to protect BHL data quality.
Qianqian Hiris Gu ↗
Biological occurrence data from historic scientific correspondence: enabling automated approaches for structured data extraction from body text
A Darwin Core-aligned annotated corpus from MfN journals provides a reference for evaluating automated occurrence extraction from historical texts.
Christian Bölling ↗
Multimodality for Knowledge Extraction from Historical Entomology Literature
Aligning text and illustrations in a multimodal knowledge graph can generate evidence-grounded captions and recover knowledge lost by text-only extraction.
Jana Hoffmann, Sefika Efeoglu ↗
Reconsidering the original vision for a global biodiversity information facility
GBIF achieved its 1999 vision only partially and through a network of collaborating infrastructures; the community must decide whether that is enough and articulate the next vision.
Joe Miller ↗
Rediscovering Archives - Using LLM's for the Transcription and Extraction of Ecological Data from Historical Archives
Off-the-shelf multimodal LLMs can transcribe varied handwritten archives well, with closed models more robust on structured tables.
Phoebe Santos ↗
How Literature Services can Support & Benefit from Biodiversity Publication and Data Standards?
Standards-based concept annotation of biodiversity literature enables focused, evidence-grounded question answering that is more precise than general chatbots.
Patrick Ruch ↗
Literature triage to support Island Biodiversity Monitoring
Classifier-based literature triage matches human curators at a fraction of LLM cost and is worth investing in for recurring curation tasks.
Patrick Ruch ↗
Literature usage in IPBES - Past, Present and Future with Libroscope
Open, scripted literature search via OpenAlex transformed IPBES practice, and the Libroscope could close remaining gaps by making paper contents and grey literature searchable.
Rainer M Krug ↗
Mobilising and integrating extracted information from the literature into a biodiversity data ecosystem - the BIOfid approach
BIOfid mobilises Central European biodiversity literature with NLP, ontologies and standards, because raw LLM extraction still hallucinates identifiers and needs verification.
Gerwin Kasperek ↗
Realtime Access to Data in new Taxonomic Publications
Plazi can deliver interlinked, machine-actionable data from new taxonomic papers within hours, and XML-first publishing removes most of the costly effort.
Guido Sautter ↗
Species trait literature mining to support integrative biodiversity distribution modelling
Demonstrating research use cases like trait-informed distribution modelling is key to persuading funders to support literature mobilisation infrastructure.
Robert Waterhouse ↗
The future is the future - thought about access to data in publications
Liberating data from literature works, but the future lies in use-case-driven, digital-first publishing that also targets the undescribed majority of species.
Donat Agosti ↗
Contained Agentic Workflow for Literature Analysis and Data Extraction in Museum Collections
Contained, local agentic RAG over a literature corpus is feasible on a laptop or shared edge machine, avoiding the security worries of cloud agents in museums.
Steen Dupont ↗
Curation through citation: using AI and a knowledge graph to curate DNA barcodes
Linking barcodes, specimens and literature in a knowledge graph could fix many BOLD records, if we can reliably extract the connecting codes from papers.
Roderic Page ↗
Extracting Structured Biodiversity Knowledge from Literature with AI-Assisted Workflows
A hybrid text-layer/OCR tool can free tables from grey-literature PDFs into CSV with about 83% precision, but it misses a third of tables, especially narrow ones.
Yağmur Güleç ↗
LLM-Based Pipeline for Extracting Nomenclatural Acts from Taxonomic Literature
A grounded, schema-constrained two-pass LLM pipeline extracts IPNI-ready nomenclatural data well once treatments are found; finding treatment boundaries is the bottleneck.
Ishaipiriyan Karunakularatnam ↗
The Planetary Knowledge Base: an infinite solutions engine converting data into action for nature
The NHM is building the Planetary Knowledge Base in phases, with provenance, attribution and benefit-sharing designed in, to move biodiversity data to evidence to action at scale.
Vincent Smith ↗