talk · Tuesday 22 September · SAL A

HerbAudit: Validating AI-Driven Herbarium Transcriptions

Dilara Ağacık · AI for Biodiversity Data

Recording time 2:57:07–3:11:07Open on Vimeo ↗

The short versionHerbAudit shows AI herbarium transcription reaching ~94% accuracy and provides a fair, field-aware way to benchmark it.

Overview

What this was about

Dilara Ağacık presented HerbAudit, a tool that both transcribes herbarium labels with a single multimodal AI call (after LeafMachine2 label detection) into Darwin Core, and evaluates transcriptions against manual annotations or GBIF records using field-specific scoring that avoids wrongly penalising AI: taxonomic synonym checks, order-independent collector matching, date-distance scoring, ISO country normalisation and transliteration, word-based locality matching and Haversine distance for coordinates. On a manually annotated benchmark of 90 specimens from 63 institutions, AI transcription averaged 94% accuracy; a Gemini 'Lite' model (ASR 'Gemini 3.5 Lite') offered the best accuracy-to-cost trade-off, the local Gemma model came close at zero API cost, and HerbAudit outperformed other tools with full extraction coverage. Handwritten labels scored lowest and most variably.

Why it matters. The community lacked a clear benchmark for AI transcription; field-aware evaluation and evidence that local models perform well help institutions adopt AI responsibly.

Key ideas

In the room

  • Manual databasing at NYBG averages 10 specimens per hour (130,000 staff hours for 1.3 million specimens).
  • GBIF records can be incomplete or wrong and are an imperfect ground truth.
  • Single-call transcription (no separate OCR) reduces time and cost and logs metadata such as model, tokens and cost.
  • Benchmark: 90 specimens, 63 institutions, handwritten/typed/mixed categories, annotated in Label Studio.
  • Interactive HTML reports show per-field and per-specimen scores; synonyms and doubtful taxonomy are excluded from scoring and flagged.
  • Error sources: hard-to-read cursive surnames, low resolution, crossed-out redeterminations and multiple labels.
Jump in

Notable moments

In their words

Transcript

Automatically generated captions can contain mistakes, especially in names and technical terms. Times are relative to the room recording.

Read the transcript ↓
Loading transcript…