Closing Gaps in Specimen Metadata: A Collaborative Annotation and Feedback Platform for Natural History Collections
Szu-Hsien Lee · Responsible AI, Open Digital Curation, and Round-Tripping for Biodiversity Data
The short versionA national collaborative annotation platform with AI-assisted transcription can close metadata gaps across many Taiwanese collections.
What this was about
Szu-Hsien Lee showed GBIF data trends indicating that specimen records increasingly lag observation records in completeness, because imaging is fast but transcription needs experts (including for Western and Japanese-era labels in Taiwan). After using students and AI within one institution, he built a system over the Taiwan Biodiversity Information Alliance (TBIA) aggregate (about 2 million specimen records) where anyone can fill missing terms with AI-assisted transcription and human correction, storing both for later evaluation. He also demonstrated collector trip views and links to historical botanical literature, and aims to return corrections to providers, arguing 'an untranscribed image is not a failure, it's an invitation'.
Why it matters. Addresses the widening completeness gap of specimen data and offers a model for institutions lacking digitisation resources.
In the room
- GBIF data trends show specimen records growing but lacking completeness (dates, identifications, coordinates), with the gap widening globally and in Taiwan.
- Imaging takes seconds; transcription of old handwriting (including Japanese-era labels) needs trained experts and funding.
- The museum's collection management system is shared with other Taiwanese institutions lacking resources; images without metadata are effectively hidden.
- College teachers assigned transcription as student homework; later AI was used, but this only helped one institution slowly.
- TBIA aggregates biodiversity data from institutions and government in Taiwan (also published to GBIF); about 2 million natural history specimen records.
- Flags: about 30% missing identification (often name-matching issues, resolved via the TaiCOL checklist), over 65% missing coordinates (often with verbatim locality), and missing media is the biggest barrier.
- Platform integrates an AI API for transcription; AI handles English handwriting well but struggles with Chinese and Japanese, so humans correct; both versions are stored for later evaluation.
Notable moments
Transcript
Automatically generated captions can contain mistakes, especially in names and technical terms. Times are relative to the room recording.


