A conference theme

LLM-based information extraction

Using large language models to extract, structure and enrich biodiversity information from text and records, including prompt and schema design.

20 talks & discussions
Across rooms and sessions

Explore the conversation.

talk · Tuesday 22 September · SAL A

A closer look at steps needed to extract research-ready biodiversity data from legacy literature

Legacy literature data can only be unlocked by combining better OCR/NLP with taxonomic rigour that accounts for synonymy, shifting concepts, coreference and mixed languages.

Gerwin Kasperek ↗
talk · Tuesday 22 September · SAL A

AI-assisted georeferencing a posteriori of herbarium specimens: a case study on Italian herbaria

Combining LLM parsing and LLM judging with gazetteers more than doubles correct georeferences for historical herbarium localities compared with Google Maps, even with a local model.

Matteo Conti ↗
talk · Tuesday 22 September · SAL A

BHL and the Planetary Knowledge Base: Structuring Biodiversity Literature for the Next Generation

NHM is re-OCRing all of BHL with modern models and extracting entities into a knowledge graph, with confidence gating to protect BHL data quality.

Qianqian Hiris Gu ↗
talk · Tuesday 22 September · SAL A

Biological occurrence data from historic scientific correspondence: enabling automated approaches for structured data extraction from body text

A Darwin Core-aligned annotated corpus from MfN journals provides a reference for evaluating automated occurrence extraction from historical texts.

Christian Bölling ↗
talk · Tuesday 22 September · SAL B

Digital curation across the biodiversity data lifecycle, at the node level (SiB Colombia), aggregator level (GBIF) and how to incorporate user feedback

Curation feedback works at every stage of the data lifecycle, and about a third or more of reported issues do get fixed; AI is used cautiously to summarise feedback.

Esteban Marentes Herrera ↗
talk · Tuesday 22 September · ODIN

From citizen observations to long-term ecological knowledge: integrating iNaturalist records into LTER framework through automated enrichment

Rather than forcing citizens to produce structured data, automated enrichment elevates iNaturalist records into research-ready data for eLTER sites.

Alessandro Oggioni ↗
talk · Tuesday 22 September · SAL B

From Images to Structured Data: A Scalable AI Workflow for Natural History Museums

A simple 'just do it' AI transcription pipeline with schema prompts, an intent step and full provenance can make label digitisation scale.

Anne Koivunen ↗
talk · Tuesday 22 September · SAL B

Geonomia: Metadata and Community for Georeferencing

Clustering specimens into collecting trips makes georeferencing and AI enrichment more efficient and should be done as a community effort.

Nicky Nicolson ↗
talk · Tuesday 22 September · SAL A

Multimodality for Knowledge Extraction from Historical Entomology Literature

Aligning text and illustrations in a multimodal knowledge graph can generate evidence-grounded captions and recover knowledge lost by text-only extraction.

Jana Hoffmann, Sefika Efeoglu ↗
talk · Tuesday 22 September · SAL B

Rediscovering Archives - Using LLM's for the Transcription and Extraction of Ecological Data from Historical Archives

Off-the-shelf multimodal LLMs can transcribe varied handwritten archives well, with closed models more robust on structured tables.

Phoebe Santos ↗
discussion · Tuesday 22 September · SAL A

SYM08B discussion: extracting geography and articles from BHL

Article-level discoverability and geographic extraction are key BHL gaps that AI may now make tractable, but resourcing remains the constraint.

Nicole Kearney, Roderic Page ↗
talk · Thursday 24 September · SAL B

AI Ready Standards with Croissant for Type Specimens Catalog Datasets

Combining LLM extraction, Darwin Core and Croissant metadata offers a route from historical specimen catalogues to FAIR, ML-ready datasets.

Sefika Efeoglu ↗
talk · Thursday 24 September · SAL C

Causal Mosaic Schema: Encoding Causal Claims in Ecological Literature as Machine-Actionable Knowledge Graphs

Encoding how and how strongly the literature claims causation, not just what causes what, can link siloed ecological knowledge into decision-support tools for practitioners.

Tim Alamenciak ↗
talk · Thursday 24 September · FORUM

Literature triage to support Island Biodiversity Monitoring

Classifier-based literature triage matches human curators at a fraction of LLM cost and is worth investing in for recurring curation tasks.

Patrick Ruch ↗
talk · Thursday 24 September · SAL C

On-Device AI for Data Cleaning, Standardisation, and Exploration in Collections Management

Local LLM hardware can clean and enrich millions of legacy collection records at predictable cost, but validation of the outputs is the unsolved problem.

Jack Hollister, Unidentified co-presenter ↗
talk · Friday 25 September · SAL A

LLM-Based Pipeline for Extracting Nomenclatural Acts from Taxonomic Literature

A grounded, schema-constrained two-pass LLM pipeline extracts IPNI-ready nomenclatural data well once treatments are found; finding treatment boundaries is the bottleneck.

Ishaipiriyan Karunakularatnam ↗
talk · Friday 25 September · SAL A

Reliable LLM-assisted curation of ecological survey data: a deployed agent bridging field collection and data curation

An LLM agent that proposes findings for curator approval, with confirmed patterns turned into deterministic rules, catches plausibility errors that schema validation misses.

Andrew Tokmakoff ↗
talk · Friday 25 September · SAL A

SCRIBE: Structured Collection Record Interpretation and Bio-entity Extraction

SCRIBE aims to replace many collection-specific extraction workflows with one LLM and computer-vision platform that turns any uploaded record images into structured data.

Arianna Salili-James ↗
talk · Friday 25 September · SAL B

Structured identification keys complementing AI: closing the corpus bottleneck with AI-assisted digitisation

AI and identification keys are complementary, and LLM-assisted digitisation of printed literature now makes it feasible to build structured keys at scale for expert review.

Wouter Koch ↗
talk · Friday 25 September · SAL A

Using Large Language Models to enhance biodiversity knowledge extraction: From taxonomic treatments to graph-based knowledge systems

Taxonomic treatments hold underused habitat knowledge that LLMs and a provenance-preserving knowledge graph can structure and link to ecoregions, helping map data-poor species.

Ian Ondo ↗