Biological occurrence data from historic scientific correspondence: enabling automated approaches for structured data extraction from body text
Christian Bölling · BHL at 20: A New Chapter for Biodiversity Literature and Data - Transforming BHL: Data Extraction, Semantic Navigation and Future Interfaces
The short versionA Darwin Core-aligned annotated corpus from MfN journals provides a reference for evaluating automated occurrence extraction from historical texts.
What this was about
Christian Bölling presented a Museum für Naturkunde Berlin project building an annotated reference corpus of historical scientific correspondence and articles from MfN journals (Deutsche Entomologische Zeitschrift, Zoosystematics and Evolution, Fossil Record and predecessors), to evaluate whether machine learning can reliably extract occurrence data at scale. A tag set (taxon, place, date, habitat, quantity, person, material, plus org/ID) aligned to Darwin Core annotates occurrences anchored at the taxon name; the first campaign found nearly 1,600 occurrence reports in 165 articles using INCEpTION. Places were matched to Wikidata (1,215 of 1,333) and taxa to GBIF IDs (1,041 of 1,154), and data serialised in RDF. He highlighted difficulties with document retrieval, erratic publisher APIs, annotation consistency and historical place names, and suggested BHL add entity recognition for persons and places and richer journal relationship metadata.
Why it matters. Legacy reports contain largely untapped occurrence data; a gold-standard corpus is needed before AI extraction can be trusted.
In the room
- Historical society meeting reports (e.g. Berlin Entomological Society 1915) detail dozens of observations.
- A valid occurrence requires a taxon name plus place and/or date.
- Places appear in almost all instances; dates, persons, quantities and materials in about a third; habitats and specimen IDs rarely.
- 681 distinct place entities and 741 distinct taxa were identified.
- Full-text coverage for some journals was much lower than articles published; new OCR was needed for many pre-1970 articles.
Notable moments
Transcript
Automatically generated captions can contain mistakes, especially in names and technical terms. Times are relative to the room recording.


