talk · Tuesday 22 September · SAL A

Biological occurrence data from historic scientific correspondence: enabling automated approaches for structured data extraction from body text

Christian Bölling · BHL at 20: A New Chapter for Biodiversity Literature and Data - Transforming BHL: Data Extraction, Semantic Navigation and Future Interfaces

Recording time 7:36:52–7:47:08Open on Vimeo ↗

The short versionA Darwin Core-aligned annotated corpus from MfN journals provides a reference for evaluating automated occurrence extraction from historical texts.

Overview

What this was about

Christian Bölling presented a Museum für Naturkunde Berlin project building an annotated reference corpus of historical scientific correspondence and articles from MfN journals (Deutsche Entomologische Zeitschrift, Zoosystematics and Evolution, Fossil Record and predecessors), to evaluate whether machine learning can reliably extract occurrence data at scale. A tag set (taxon, place, date, habitat, quantity, person, material, plus org/ID) aligned to Darwin Core annotates occurrences anchored at the taxon name; the first campaign found nearly 1,600 occurrence reports in 165 articles using INCEpTION. Places were matched to Wikidata (1,215 of 1,333) and taxa to GBIF IDs (1,041 of 1,154), and data serialised in RDF. He highlighted difficulties with document retrieval, erratic publisher APIs, annotation consistency and historical place names, and suggested BHL add entity recognition for persons and places and richer journal relationship metadata.

Why it matters. Legacy reports contain largely untapped occurrence data; a gold-standard corpus is needed before AI extraction can be trusted.

Key ideas

In the room

  • Historical society meeting reports (e.g. Berlin Entomological Society 1915) detail dozens of observations.
  • A valid occurrence requires a taxon name plus place and/or date.
  • Places appear in almost all instances; dates, persons, quantities and materials in about a third; habitats and specimen IDs rarely.
  • 681 distinct place entities and 741 distinct taxa were identified.
  • Full-text coverage for some journals was much lower than articles published; new OCR was needed for many pre-1970 articles.
Jump in

Notable moments

In their words

Transcript

Automatically generated captions can contain mistakes, especially in names and technical terms. Times are relative to the room recording.

Read the transcript ↓
Loading transcript…