Explore the conversation.

A closer look at steps needed to extract research-ready biodiversity data from legacy literature
Legacy literature data can only be unlocked by combining better OCR/NLP with taxonomic rigour that accounts for synonymy, shifting concepts, coreference and mixed languages.
Gerwin Kasperek ↗
AI-assisted georeferencing a posteriori of herbarium specimens: a case study on Italian herbaria
Combining LLM parsing and LLM judging with gazetteers more than doubles correct georeferences for historical herbarium localities compared with Google Maps, even with a local model.
Matteo Conti ↗
BHL and the Planetary Knowledge Base: Structuring Biodiversity Literature for the Next Generation
NHM is re-OCRing all of BHL with modern models and extracting entities into a knowledge graph, with confidence gating to protect BHL data quality.
Qianqian Hiris Gu ↗
Biological occurrence data from historic scientific correspondence: enabling automated approaches for structured data extraction from body text
A Darwin Core-aligned annotated corpus from MfN journals provides a reference for evaluating automated occurrence extraction from historical texts.
Christian Bölling ↗
Digital curation across the biodiversity data lifecycle, at the node level (SiB Colombia), aggregator level (GBIF) and how to incorporate user feedback
Curation feedback works at every stage of the data lifecycle, and about a third or more of reported issues do get fixed; AI is used cautiously to summarise feedback.
Esteban Marentes Herrera ↗
From citizen observations to long-term ecological knowledge: integrating iNaturalist records into LTER framework through automated enrichment
Rather than forcing citizens to produce structured data, automated enrichment elevates iNaturalist records into research-ready data for eLTER sites.
Alessandro Oggioni ↗
From Images to Structured Data: A Scalable AI Workflow for Natural History Museums
A simple 'just do it' AI transcription pipeline with schema prompts, an intent step and full provenance can make label digitisation scale.
Anne Koivunen ↗
Geonomia: Metadata and Community for Georeferencing
Clustering specimens into collecting trips makes georeferencing and AI enrichment more efficient and should be done as a community effort.
Nicky Nicolson ↗
Multimodality for Knowledge Extraction from Historical Entomology Literature
Aligning text and illustrations in a multimodal knowledge graph can generate evidence-grounded captions and recover knowledge lost by text-only extraction.
Jana Hoffmann, Sefika Efeoglu ↗
Rediscovering Archives - Using LLM's for the Transcription and Extraction of Ecological Data from Historical Archives
Off-the-shelf multimodal LLMs can transcribe varied handwritten archives well, with closed models more robust on structured tables.
Phoebe Santos ↗
SYM08B discussion: extracting geography and articles from BHL
Article-level discoverability and geographic extraction are key BHL gaps that AI may now make tractable, but resourcing remains the constraint.
Nicole Kearney, Roderic Page ↗
AI Ready Standards with Croissant for Type Specimens Catalog Datasets
Combining LLM extraction, Darwin Core and Croissant metadata offers a route from historical specimen catalogues to FAIR, ML-ready datasets.
Sefika Efeoglu ↗
Causal Mosaic Schema: Encoding Causal Claims in Ecological Literature as Machine-Actionable Knowledge Graphs
Encoding how and how strongly the literature claims causation, not just what causes what, can link siloed ecological knowledge into decision-support tools for practitioners.
Tim Alamenciak ↗
Literature triage to support Island Biodiversity Monitoring
Classifier-based literature triage matches human curators at a fraction of LLM cost and is worth investing in for recurring curation tasks.
Patrick Ruch ↗
On-Device AI for Data Cleaning, Standardisation, and Exploration in Collections Management
Local LLM hardware can clean and enrich millions of legacy collection records at predictable cost, but validation of the outputs is the unsolved problem.
Jack Hollister, Unidentified co-presenter ↗
LLM-Based Pipeline for Extracting Nomenclatural Acts from Taxonomic Literature
A grounded, schema-constrained two-pass LLM pipeline extracts IPNI-ready nomenclatural data well once treatments are found; finding treatment boundaries is the bottleneck.
Ishaipiriyan Karunakularatnam ↗
Reliable LLM-assisted curation of ecological survey data: a deployed agent bridging field collection and data curation
An LLM agent that proposes findings for curator approval, with confirmed patterns turned into deterministic rules, catches plausibility errors that schema validation misses.
Andrew Tokmakoff ↗
SCRIBE: Structured Collection Record Interpretation and Bio-entity Extraction
SCRIBE aims to replace many collection-specific extraction workflows with one LLM and computer-vision platform that turns any uploaded record images into structured data.
Arianna Salili-James ↗
Structured identification keys complementing AI: closing the corpus bottleneck with AI-assisted digitisation
AI and identification keys are complementary, and LLM-assisted digitisation of printed literature now makes it feasible to build structured keys at scale for expert review.
Wouter Koch ↗
Using Large Language Models to enhance biodiversity knowledge extraction: From taxonomic treatments to graph-based knowledge systems
Taxonomic treatments hold underused habitat knowledge that LLMs and a provenance-preserving knowledge graph can structure and link to ecoregions, helping map data-poor species.
Ian Ondo ↗