Structured identification keys complementing AI: closing the corpus bottleneck with AI-assisted digitisation
Wouter Koch · May the Data Be Structured: Linking Descriptions, Identification and AI
The short versionAI and identification keys are complementary, and LLM-assisted digitisation of printed literature now makes it feasible to build structured keys at scale for expert review.
What this was about
Koch argued that image-recognition AI and structured multi-access identification keys complement each other. In his pipeline an image model gives candidate species, which then filter a Clavis-format key so the user answers only a few discriminating questions. He also showed a prototype in which a vision-language model checks which key characters are visible in a photo and explains the identification. To remove the key-building bottleneck, an agentic LLM pipeline extracts characters from printed literature (e.g., Norwegian mammal books) into draft keys for around $10-30 each, which experts then review.
Why it matters. Keys add explanation and education that image classifiers lack, and they ground AI in domain knowledge that is disappearing as taxonomists retire. Cheap draft keys change the expert's role from building keys to reviewing them.
In the room
- Every occurrence point depends on an identification, so identification underpins most biodiversity data.
- The Norwegian image-recognition app is kept low-threshold (no sign-up) and shows uncertainty, e.g. warning users not to eat mushrooms identified by AI; but images cannot capture every diagnostic feature and give no learning.
- Digital multi-access keys avoid getting stuck on an unanswerable question. Keys are authored in the Clavis JSON format, and Koch argued that AI makes mapping between formats increasingly trivial.
- Pipeline 1: photo, then image AI species candidates, then those candidates filter the key, so the user answers two or three species-specific questions. Some states are already eliminated, which reduces user error. A working dragonfly demo is integrated with the national app API.
- Pipeline 2 (not yet scalable): candidate list, key facts and photo are passed to a multimodal VLM, which reports which characters it can see and generates an explanation that includes diagnostic features not visible in the photo.
- An agentic pipeline extracts facts only from supplied sources (e.g., rodent traits from Norwegian mammal books) and collates them into a key, possibly one that never existed as a key before.
- Draft keys cost under $10 for a single source and about $30 for multiple sources with Claude Opus. Experts now review a working key in an interface instead of starting from an empty spreadsheet.
Notable moments
Transcript
Automatically generated captions can contain mistakes, especially in names and technical terms. Times are relative to the room recording.


