A conference theme

Label transcription and OCR

Transcribing specimen labels, cards, handwritten records and printed text with OCR, handwriting recognition and multimodal LLMs, and evaluating transcription quality.

16 talks & discussions
Across rooms and sessions

Explore the conversation.

talk · Tuesday 22 September · SAL A

A closer look at steps needed to extract research-ready biodiversity data from legacy literature

Legacy literature data can only be unlocked by combining better OCR/NLP with taxonomic rigour that accounts for synonymy, shifting concepts, coreference and mixed languages.

Gerwin Kasperek ↗
talk · Tuesday 22 September · SAL A

BHL and the Planetary Knowledge Base: Structuring Biodiversity Literature for the Next Generation

NHM is re-OCRing all of BHL with modern models and extracting entities into a knowledge graph, with confidence gating to protect BHL data quality.

Qianqian Hiris Gu ↗
talk · Tuesday 22 September · SAL B

Closing Gaps in Specimen Metadata: A Collaborative Annotation and Feedback Platform for Natural History Collections

A national collaborative annotation platform with AI-assisted transcription can close metadata gaps across many Taiwanese collections.

Szu-Hsien Lee ↗
talk · Tuesday 22 September · SAL A

For the Birds: Adapting Archives and Extracting Data for the Biodiversity Heritage Library

USF is strengthening BHL's infrastructure via a redundant Macaw ingest instance while working to bring vulnerable Florida environmental archives and their data into BHL.

Amanda Boczar, Sydney Jordan ↗
talk · Tuesday 22 September · SAL B

From Images to Structured Data: A Scalable AI Workflow for Natural History Museums

A simple 'just do it' AI transcription pipeline with schema prompts, an intent step and full provenance can make label digitisation scale.

Anne Koivunen ↗
talk · Tuesday 22 September · SAL A

Future interfaces for BHL

Modern AI tools make new BHL interfaces feasible now (better OCR, image and map search, chatbots); the community must decide which are worth building.

Roderic Page ↗
talk · Tuesday 22 September · SAL A

HerbAudit: Validating AI-Driven Herbarium Transcriptions

HerbAudit shows AI herbarium transcription reaching ~94% accuracy and provides a fair, field-aware way to benchmark it.

Dilara Ağacık ↗
talk · Tuesday 22 September · SAL B

KakraCards: An AI-Assisted Pipeline for Liberating Six Decades of Seabird Heritage Data

Multi-model consensus with human adjudication builds ground truth and picks the best LLM for transcribing standardised historical cards.

Kristjan Adojaan ↗
talk · Tuesday 22 September · SAL C

Pinned insect digitization conveyor: image and data transcription workflows

An industrial conveyor workflow with QR-encoded metadata and automated skeletal records makes mass digitisation of pinned insects feasible.

Jessica Bird, Sylvia Orli ↗
talk · Tuesday 22 September · SAL B

Rediscovering Archives - Using LLM's for the Transcription and Extraction of Ecological Data from Historical Archives

Off-the-shelf multimodal LLMs can transcribe varied handwritten archives well, with closed models more robust on structured tables.

Phoebe Santos ↗
talk · Tuesday 22 September · SAL B

Robots Rapidly Writing Records: Machine Annotations at Scale with DiSSCo

DiSSCo can now run adapted machine annotation services over whole datasets, pointing to an image-to-verified-data digitisation pipeline.

Soulaine Theocharides ↗
talk · Thursday 24 September · SAL A

A framework for aligning collections data with the global data ecosystem from data creation to extension in the Smithsonian National Museum of Natural History Department of Paleobiology

Making paleo collections work in the global data ecosystem needs an iterative, full life-cycle framework, not one-off digitisation.

Holly Little ↗
talk · Friday 25 September · SAL A

Automated data transcription and Darwin Core field itemization from insect specimen labels using machine learning

Labellum goes beyond label OCR to itemise insect-label text into Darwin Core fields, with high accuracy for most fields but problems with catalogue numbers.

Torsten Dikow ↗
talk · Friday 25 September · SAL A

Extracting Structured Biodiversity Knowledge from Literature with AI-Assisted Workflows

A hybrid text-layer/OCR tool can free tables from grey-literature PDFs into CSV with about 83% precision, but it misses a third of tables, especially narrow ones.

Yağmur Güleç ↗
talk · Friday 25 September · SAL A

From Historical Oology Cards to Structured Biodiversity Data: Multimodal LLM Transcription, Uncertainty Metrics, and Human Review

Token-level hesitation from log-probabilities does not tell you which transcriptions are wrong, but it reliably predicts where human review effort will go.

Grete Pasch ↗
talk · Friday 25 September · SAL A

SCRIBE: Structured Collection Record Interpretation and Bio-entity Extraction

SCRIBE aims to replace many collection-specific extraction workflows with one LLM and computer-vision platform that turns any uploaded record images into structured data.

Arianna Salili-James ↗