talk · Thursday 24 September · SAL C

Mobilising and integrating extracted information from the literature into a biodiversity data ecosystem - the BIOfid approach

Gerwin Kasperek · From Mobilizing Data to AI-Ready Knowledge: Infrastructure for Multimodal Biodiversity Data

Recording time 6:17:08–6:30:40Open on Vimeo ↗

The short versionBIOfid mobilises Central European biodiversity literature with NLP, ontologies and standards, because raw LLM extraction still hallucinates identifiers and needs verification.

Overview

What this was about

Gerwin Kasperek argued that literature is the only biodiversity data source combining historical depth with high volume, but its data are locked in multilingual, abbreviation-heavy, sometimes poorly printed text. Using an 18th-century German passage about an orchid near Frankfurt, he showed the occurrence and trait records that could be extracted. He then showed a current LLM getting taxon and habitat right but hallucinating GBIF and Wikidata IDs, even after prompt improvements, so outputs must be verified against authoritative sources. He presented BIOfid, the DFG-funded specialised information service run by Senckenberg and Goethe University Frankfurt, with its annotated corpus, Unified Corpus Explorer, NLP pipelines, ontologies and plans for multimodal extraction into RDF/OWL knowledge graphs.

Why it matters. Centuries of occurrence and trait data exist only in print. Standard-based, verifiable extraction pipelines are needed before such data can safely feed knowledge graphs and AI.

Key ideas

In the room

  • Collections data have depth but limited volume, and citizen science has volume but only recent years; literature combines both but is the hardest to access.
  • Challenges: natural language, non-English texts (German, Latin, Hungarian), insider abbreviations, domain conventions such as indented keys, poor print quality.
  • Mobilisation means converting analog, unstructured information into FAIR digital formats, not just digitisation.
  • Example: an occurrence of Coeloglossum viride near Eschenheimer Tor, Frankfurt, May 1729, in a meadow, plus a leaf-width trait record, extracted from one sentence.
  • An LLM extracted taxon and habitat correctly but hallucinated GBIF and Wikidata IDs, still wrong after a stricter prompt; GBIF and World Flora Online also disagree on the accepted name (Coeloglossum viride vs Dactylorhiza viridis).
  • BIOfid aims: mobilise data from Central European literature, provide text-mining tools, digitise 20th-century literature; a collaboration of Senckenberg, Goethe University text technology and the Senckenberg Library, funded by DFG.
  • Uses the generic standards UIMA, OWL and RDF/SPARQL and the domain standards Darwin Core and GBIF/Catalogue of Life backbones. Reusable components are DUUI, the Unified Corpus Explorer (UCE) and ontologies (e.g. Chilopoda anatomy, EnvO extensions and translations).
Jump in

Notable moments

In their words

Transcript

Automatically generated captions can contain mistakes, especially in names and technical terms. Times are relative to the room recording.

Read the transcript ↓
Loading transcript…