From Historical Oology Cards to Structured Biodiversity Data: Multimodal LLM Transcription, Uncertainty Metrics, and Human Review
Grete Pasch · Large Language Models for Biodiversity Data Discovery, Integration, and Curation. Part A: AI- enabled knowledge creation and extraction
The short versionToken-level hesitation from log-probabilities does not tell you which transcriptions are wrong, but it reliably predicts where human review effort will go.
What this was about
Grete Pasch transcribed historical oology (egg-set) cards with GPT-4o and used token log-probabilities to build a 'hesitation' score. The score does not flag every error, but it predicts how much human review a card needs. Cards were sorted into light, standard and expert review lanes; the expert lane held about a fifth of the cards but took half to two-thirds of the correction time, on both the Cornell set and 2,303 Western Foundation cards. Newer models (Gemini 3.1 Pro, Opus 5) are more accurate but do not expose log-probs; where two of them agree they are right about 997 times in 1,000, but that still does not certify a transcription.
Why it matters. Most LLM transcription work stops at accuracy. This talk offers a practical way to triage expert review and estimate curation effort for millions of archival cards, and warns that the best models no longer expose the signals needed for it.
In the room
- Nest descriptions on historical egg cards record changes in nest materials (e.g. plastic in modern Baltimore oriole nests), but they are rarely transcribed.
- GPT-4o transcribed 193 of 262 cards exactly (about three in four). Typed cards were nearly perfect; handwritten cards were about two in three exact.
- With logprobs enabled, the model reports its confidence in each token and its top alternatives. On a gull card it wrote '7,000' at 58% where the correct '2,000' was the runner-up at 27%.
- A token counts as 'hesitant' when the top choice is less than 2.7 times as likely as the runner-up; the card hesitation score is the number of hesitant tokens.
- Some cards with no hesitation still had errors (confident errors), so hesitation marks uncertainty, not correctness.
- Three lanes (0-1, 2-4, 5+ hesitant tokens): median correction time was under a second in the light and standard lanes and 11 seconds in the expert lane.
- The pattern held on 2,303 harder handwritten cards from the Western Foundation of Vertebrate Zoology.
Notable moments
Transcript
Automatically generated captions can contain mistakes, especially in names and technical terms. Times are relative to the room recording.


