talk · Tuesday 22 September · SAL A

BHL and the Planetary Knowledge Base: Structuring Biodiversity Literature for the Next Generation

Qianqian Hiris Gu · BHL at 20: A New Chapter for Biodiversity Literature and Data - Reimagining BHL: Foundations, Partnerships and Future Infrastructure

Recording time 6:22:42–6:41:12Open on Vimeo ↗

The short versionNHM is re-OCRing all of BHL with modern models and extracting entities into a knowledge graph, with confidence gating to protect BHL data quality.

Overview

What this was about

Qianqian Hiris Gu described NHM London's work making BHL machine-readable and connecting it to the NHM's Planetary Knowledge Base. Building on Neil Brummitt's decade of trait extraction from English-language BHL literature (about 700,000 trait terms across 81 standardised traits), an AWS prototyping sprint of four weeks on under 40,000 pages of a botanical journal combined modern OCR with computer vision for article title-page detection (~98%, ~100% with a table of contents), entity extraction (taxa, people, localities, institutions, habitats; F1 ~96%) into a knowledge graph and GraphRAG search. Benchmarking with BHL-provided ground truth showed printed text reaching human-level transcription, and word-level confidence from OCR models can gate what goes back into BHL. The work is now in production: re-OCR of the whole corpus by around January 2027, an open gold-standard dataset in early 2027, and an Explorer interface in Q2 2027.

Why it matters. Turning BHL's 64 million pages into structured, linkable data could unlock huge amounts of trait, taxon and specimen information for research and infrastructure links.

Key ideas

In the room

  • Four pages of a 2002 publication contained 142 data points; BHL has about 64 million pages.
  • Previous trait extraction covered only English literature and 17% of accepted flowering plants.
  • Earlier pipelines needed a separate model per task; LLMs/VLMs allow more in one pass.
  • Article segmentation via title page detection enables assigning DOIs and references.
  • Word-level OCR confidence creates a barrier against ingesting hallucinated text into BHL.
  • Production also extracts specimen identifiers to link back to GBIF and stratigraphy data.
Jump in

Notable moments

In their words

Transcript

Automatically generated captions can contain mistakes, especially in names and technical terms. Times are relative to the room recording.

Read the transcript ↓
Loading transcript…