A conference theme

Biodiversity literature mining

Extracting occurrences, traits, names and other data from biodiversity literature and grey literature with text mining and NLP.

19 talks & discussions
Across rooms and sessions

Explore the conversation.

talk · Monday 21 September · Aula

Building a New Foundation for the Biodiversity Heritage Library

BHL has survived losing its Smithsonian home by becoming an independent, globally governed consortium, but its future is still not secure and it needs community support.

Nicole Kearney ↗
talk · Tuesday 22 September · SAL A

A closer look at steps needed to extract research-ready biodiversity data from legacy literature

Legacy literature data can only be unlocked by combining better OCR/NLP with taxonomic rigour that accounts for synonymy, shifting concepts, coreference and mixed languages.

Gerwin Kasperek ↗
talk · Tuesday 22 September · SAL A

BHL and the Planetary Knowledge Base: Structuring Biodiversity Literature for the Next Generation

NHM is re-OCRing all of BHL with modern models and extracting entities into a knowledge graph, with confidence gating to protect BHL data quality.

Qianqian Hiris Gu ↗
talk · Tuesday 22 September · SAL A

Biological occurrence data from historic scientific correspondence: enabling automated approaches for structured data extraction from body text

A Darwin Core-aligned annotated corpus from MfN journals provides a reference for evaluating automated occurrence extraction from historical texts.

Christian Bölling ↗
talk · Tuesday 22 September · SAL A

Multimodality for Knowledge Extraction from Historical Entomology Literature

Aligning text and illustrations in a multimodal knowledge graph can generate evidence-grounded captions and recover knowledge lost by text-only extraction.

Jana Hoffmann, Sefika Efeoglu ↗
talk · Tuesday 22 September · SAL A

Reconsidering the original vision for a global biodiversity information facility

GBIF achieved its 1999 vision only partially and through a network of collaborating infrastructures; the community must decide whether that is enough and articulate the next vision.

Joe Miller ↗
talk · Tuesday 22 September · SAL B

Rediscovering Archives - Using LLM's for the Transcription and Extraction of Ecological Data from Historical Archives

Off-the-shelf multimodal LLMs can transcribe varied handwritten archives well, with closed models more robust on structured tables.

Phoebe Santos ↗
talk · Thursday 24 September · SAL C

How Literature Services can Support & Benefit from Biodiversity Publication and Data Standards?

Standards-based concept annotation of biodiversity literature enables focused, evidence-grounded question answering that is more precise than general chatbots.

Patrick Ruch ↗
talk · Thursday 24 September · FORUM

Literature triage to support Island Biodiversity Monitoring

Classifier-based literature triage matches human curators at a fraction of LLM cost and is worth investing in for recurring curation tasks.

Patrick Ruch ↗
talk · Thursday 24 September · FORUM

Literature usage in IPBES - Past, Present and Future with Libroscope

Open, scripted literature search via OpenAlex transformed IPBES practice, and the Libroscope could close remaining gaps by making paper contents and grey literature searchable.

Rainer M Krug ↗
talk · Thursday 24 September · SAL C

Mobilising and integrating extracted information from the literature into a biodiversity data ecosystem - the BIOfid approach

BIOfid mobilises Central European biodiversity literature with NLP, ontologies and standards, because raw LLM extraction still hallucinates identifiers and needs verification.

Gerwin Kasperek ↗
talk · Thursday 24 September · FORUM

Realtime Access to Data in new Taxonomic Publications

Plazi can deliver interlinked, machine-actionable data from new taxonomic papers within hours, and XML-first publishing removes most of the costly effort.

Guido Sautter ↗
talk · Thursday 24 September · FORUM

Species trait literature mining to support integrative biodiversity distribution modelling

Demonstrating research use cases like trait-informed distribution modelling is key to persuading funders to support literature mobilisation infrastructure.

Robert Waterhouse ↗
talk · Thursday 24 September · FORUM

The future is the future - thought about access to data in publications

Liberating data from literature works, but the future lies in use-case-driven, digital-first publishing that also targets the undescribed majority of species.

Donat Agosti ↗
talk · Friday 25 September · SAL A

Contained Agentic Workflow for Literature Analysis and Data Extraction in Museum Collections

Contained, local agentic RAG over a literature corpus is feasible on a laptop or shared edge machine, avoiding the security worries of cloud agents in museums.

Steen Dupont ↗
talk · Friday 25 September · SAL A

Curation through citation: using AI and a knowledge graph to curate DNA barcodes

Linking barcodes, specimens and literature in a knowledge graph could fix many BOLD records, if we can reliably extract the connecting codes from papers.

Roderic Page ↗
talk · Friday 25 September · SAL A

Extracting Structured Biodiversity Knowledge from Literature with AI-Assisted Workflows

A hybrid text-layer/OCR tool can free tables from grey-literature PDFs into CSV with about 83% precision, but it misses a third of tables, especially narrow ones.

Yağmur Güleç ↗
talk · Friday 25 September · SAL A

LLM-Based Pipeline for Extracting Nomenclatural Acts from Taxonomic Literature

A grounded, schema-constrained two-pass LLM pipeline extracts IPNI-ready nomenclatural data well once treatments are found; finding treatment boundaries is the bottleneck.

Ishaipiriyan Karunakularatnam ↗
talk · Friday 25 September · SAL A

The Planetary Knowledge Base: an infinite solutions engine converting data into action for nature

The NHM is building the Planetary Knowledge Base in phases, with provenance, attribution and benefit-sharing designed in, to move biodiversity data to evidence to action at scale.

Vincent Smith ↗