From Images to Structured Data: A Scalable AI Workflow for Natural History Museums
Anne Koivunen · Responsible AI, Open Digital Curation, and Round-Tripping for Biodiversity Data
The short versionA simple 'just do it' AI transcription pipeline with schema prompts, an intent step and full provenance can make label digitisation scale.
What this was about
Anne Koivunen, head of digitisation at Luomus, described a pragmatic AI pipeline for tackling a backlog of 10 million undigitised specimens with fewer staff and a goal of full digitisation by 2045. Label images go to Gemini 2.5 for verbatim transcription, then again with a prompt and schema for structured data saved into the Kotka collection management system; for messy vascular plant and moss labels an 'intent' step asks the model what the writer meant. AI output never overrides human-transcribed fields, is flagged, and run metadata (model version, prompts) is stored. She stressed that winning collection staff's trust was harder than the AI itself.
Why it matters. A realistic, production-oriented example from a mid-sized museum showing both technical choices and the organisational change needed to adopt AI transcription.
In the room
- Luomus (University of Helsinki) has about 14 million specimens, about 10 million invertebrates; about 4 million digitised in ~10 years of mass digitisation.
- Strategy goal: digitise everything by 2045, requiring about 30% faster digitisation despite staff cuts.
- Scalability means variety as well as volume: taxon-specific label types, handwritten/typed, multilingual (Finnish, Swedish, Russian), corrections and differing collection traditions.
- Pipeline: images -> Gemini 2.5 verbatim transcription (1-5 hours per batch) -> post-processing -> Gemini with prompt and schema for structuring -> Kotka CMS; the schema is the most important part of the prompt.
- Complex labels get an extra 'intent' step interpreting what the writer meant, giving near-correct output even for heavily corrected labels.
- Field names were partly invented rather than strict Darwin Core terms because this gave better results.
- Verbatim text is stored in a dedicated field; structured data goes to specific fields, flagged as AI-generated, never overriding human transcription; model versions and prompts are stored.
Notable moments
Transcript
Automatically generated captions can contain mistakes, especially in names and technical terms. Times are relative to the room recording.


