talk · Tuesday 22 September · SAL B

Rediscovering Archives - Using LLM's for the Transcription and Extraction of Ecological Data from Historical Archives

Phoebe Santos · Bots, Bits, and Biodiversity

Recording time 7:31:45–7:36:55Open on Vimeo ↗

The short versionOff-the-shelf multimodal LLMs can transcribe varied handwritten archives well, with closed models more robust on structured tables.

Overview

What this was about

Phoebe Santos presented master's research (NHM, ZSL, UCL) testing whether a single LLM transcription approach generalises across different historical record types. Using British Library colonial-era India records and NHM mammal osteology table records, a closed-weight Claude Sonnet model and open-weight Qwen 2.5-VL 32B were prompted to classify page structure (text, mixed, table) and transcribe while preserving structure, with zero-, one- and few-shot and dataset-specific examples. Both did well on text; Claude was significantly better especially on tables and was robust to prompting, while Qwen improved with more and dataset-specific examples.

Why it matters. Evidence that LLM transcription is a viable, affordable route to unlocking historical ecological archives.

Key ideas

In the room

  • Goal: test generalisation of one approach across collections with different handwriting, format and structure.
  • Structure matters for meaning; models often failed to reproduce page formatting.
  • Datasets: British Library 18th-19th century British Colonial Office India records; NHM handwritten mammal osteology tables.
  • Models: Claude Sonnet (closed weights) and Qwen 2.5-VL 32B (open weights), temperature 0.7, no other changes.
  • Prompt asks the model to classify each page as text, mixed or table and apply matching structure instructions.
  • Both good on text; Claude significantly better (small absolute gap) and much better on structured records; Claude consistent across prompting; Qwen improved with more, dataset-specific examples.
  • Claude costs suggested LLM transcription may be cheaper than commercial services; limitations: only English, left-to-right scripts.
Jump in

Notable moments

In their words

Transcript

Automatically generated captions can contain mistakes, especially in names and technical terms. Times are relative to the room recording.

Read the transcript ↓
Loading transcript…