Rediscovering Archives - Using LLM's for the Transcription and Extraction of Ecological Data from Historical Archives
Phoebe Santos · Bots, Bits, and Biodiversity
The short versionOff-the-shelf multimodal LLMs can transcribe varied handwritten archives well, with closed models more robust on structured tables.
What this was about
Phoebe Santos presented master's research (NHM, ZSL, UCL) testing whether a single LLM transcription approach generalises across different historical record types. Using British Library colonial-era India records and NHM mammal osteology table records, a closed-weight Claude Sonnet model and open-weight Qwen 2.5-VL 32B were prompted to classify page structure (text, mixed, table) and transcribe while preserving structure, with zero-, one- and few-shot and dataset-specific examples. Both did well on text; Claude was significantly better especially on tables and was robust to prompting, while Qwen improved with more and dataset-specific examples.
Why it matters. Evidence that LLM transcription is a viable, affordable route to unlocking historical ecological archives.
In the room
- Goal: test generalisation of one approach across collections with different handwriting, format and structure.
- Structure matters for meaning; models often failed to reproduce page formatting.
- Datasets: British Library 18th-19th century British Colonial Office India records; NHM handwritten mammal osteology tables.
- Models: Claude Sonnet (closed weights) and Qwen 2.5-VL 32B (open weights), temperature 0.7, no other changes.
- Prompt asks the model to classify each page as text, mixed or table and apply matching structure instructions.
- Both good on text; Claude significantly better (small absolute gap) and much better on structured records; Claude consistent across prompting; Qwen improved with more, dataset-specific examples.
- Claude costs suggested LLM transcription may be cheaper than commercial services; limitations: only English, left-to-right scripts.
Notable moments
Transcript
Automatically generated captions can contain mistakes, especially in names and technical terms. Times are relative to the room recording.


