SCRIBE: Structured Collection Record Interpretation and Bio-entity Extraction
Arianna Salili-James · Large Language Models for Biodiversity Data Discovery, Integration, and Curation. Part A: AI- enabled knowledge creation and extraction
The short versionSCRIBE aims to replace many collection-specific extraction workflows with one LLM and computer-vision platform that turns any uploaded record images into structured data.
What this was about
Arianna Salili-James reviewed the NHM's earlier data-extraction projects: ALICE label imaging, video label tracking, egg-card structuring with Google Vision (over 90% accuracy), and LLM extraction from registers, bird-skin labels, herbarium sheets and double-sided cards. Each needed its own workflow and prompts. SCRIBE, funded by DCMS and a private donor, is meant to replace these with one platform where users upload images and get structured output. It clusters mixed datasets, samples each cluster to build customised prompts, runs LLM extraction and cleans the output towards Darwin Core. The project started recently (data scientist hired in August) and runs to March 2028.
Why it matters. A generic, hosted extraction service could let DiSSCo UK institutions without in-house AI capacity turn labels, cards and registers into structured data, and the output would feed the Planetary Knowledge Base.
In the room
- Previous NHM extraction projects each had separate workflows, training and prompts, which was inefficient.
- SCRIBE targets information that can be output in a structured form (e.g. locality, habitat), not free-text transcription of books.
- Users say whether an upload is homogeneous or mixed. Mixed datasets are clustered (tests with DBSCAN and PCA) before prompts are built.
- Samples from each cluster go through a quick LLM entity check to customise a base prompt, then extraction runs on the full dataset, followed by cleaning and Darwin Core formatting.
- Pre-processing includes auto-rotation and auto-cropping.
- DiSSCo UK partners will get free token allowances; others can add their own API key. Shared outputs may feed the PKB.
- Tests found handwritten and printed/typed text did equally well with LLM extraction.
Notable moments
Transcript
Automatically generated captions can contain mistakes, especially in names and technical terms. Times are relative to the room recording.


