Explore the conversation.

A closer look at steps needed to extract research-ready biodiversity data from legacy literature
Legacy literature data can only be unlocked by combining better OCR/NLP with taxonomic rigour that accounts for synonymy, shifting concepts, coreference and mixed languages.
Gerwin Kasperek ↗
BHL and the Planetary Knowledge Base: Structuring Biodiversity Literature for the Next Generation
NHM is re-OCRing all of BHL with modern models and extracting entities into a knowledge graph, with confidence gating to protect BHL data quality.
Qianqian Hiris Gu ↗
Closing Gaps in Specimen Metadata: A Collaborative Annotation and Feedback Platform for Natural History Collections
A national collaborative annotation platform with AI-assisted transcription can close metadata gaps across many Taiwanese collections.
Szu-Hsien Lee ↗
For the Birds: Adapting Archives and Extracting Data for the Biodiversity Heritage Library
USF is strengthening BHL's infrastructure via a redundant Macaw ingest instance while working to bring vulnerable Florida environmental archives and their data into BHL.
Amanda Boczar, Sydney Jordan ↗
From Images to Structured Data: A Scalable AI Workflow for Natural History Museums
A simple 'just do it' AI transcription pipeline with schema prompts, an intent step and full provenance can make label digitisation scale.
Anne Koivunen ↗
Future interfaces for BHL
Modern AI tools make new BHL interfaces feasible now (better OCR, image and map search, chatbots); the community must decide which are worth building.
Roderic Page ↗
HerbAudit: Validating AI-Driven Herbarium Transcriptions
HerbAudit shows AI herbarium transcription reaching ~94% accuracy and provides a fair, field-aware way to benchmark it.
Dilara Ağacık ↗
KakraCards: An AI-Assisted Pipeline for Liberating Six Decades of Seabird Heritage Data
Multi-model consensus with human adjudication builds ground truth and picks the best LLM for transcribing standardised historical cards.
Kristjan Adojaan ↗
Pinned insect digitization conveyor: image and data transcription workflows
An industrial conveyor workflow with QR-encoded metadata and automated skeletal records makes mass digitisation of pinned insects feasible.
Jessica Bird, Sylvia Orli ↗
Rediscovering Archives - Using LLM's for the Transcription and Extraction of Ecological Data from Historical Archives
Off-the-shelf multimodal LLMs can transcribe varied handwritten archives well, with closed models more robust on structured tables.
Phoebe Santos ↗
Robots Rapidly Writing Records: Machine Annotations at Scale with DiSSCo
DiSSCo can now run adapted machine annotation services over whole datasets, pointing to an image-to-verified-data digitisation pipeline.
Soulaine Theocharides ↗
A framework for aligning collections data with the global data ecosystem from data creation to extension in the Smithsonian National Museum of Natural History Department of Paleobiology
Making paleo collections work in the global data ecosystem needs an iterative, full life-cycle framework, not one-off digitisation.
Holly Little ↗
Automated data transcription and Darwin Core field itemization from insect specimen labels using machine learning
Labellum goes beyond label OCR to itemise insect-label text into Darwin Core fields, with high accuracy for most fields but problems with catalogue numbers.
Torsten Dikow ↗
Extracting Structured Biodiversity Knowledge from Literature with AI-Assisted Workflows
A hybrid text-layer/OCR tool can free tables from grey-literature PDFs into CSV with about 83% precision, but it misses a third of tables, especially narrow ones.
Yağmur Güleç ↗
From Historical Oology Cards to Structured Biodiversity Data: Multimodal LLM Transcription, Uncertainty Metrics, and Human Review
Token-level hesitation from log-probabilities does not tell you which transcriptions are wrong, but it reliably predicts where human review effort will go.
Grete Pasch ↗
SCRIBE: Structured Collection Record Interpretation and Bio-entity Extraction
SCRIBE aims to replace many collection-specific extraction workflows with one LLM and computer-vision platform that turns any uploaded record images into structured data.
Arianna Salili-James ↗