Mobilising and integrating extracted information from the literature into a biodiversity data ecosystem - the BIOfid approach
Gerwin Kasperek · From Mobilizing Data to AI-Ready Knowledge: Infrastructure for Multimodal Biodiversity Data
The short versionBIOfid mobilises Central European biodiversity literature with NLP, ontologies and standards, because raw LLM extraction still hallucinates identifiers and needs verification.
What this was about
Gerwin Kasperek argued that literature is the only biodiversity data source combining historical depth with high volume, but its data are locked in multilingual, abbreviation-heavy, sometimes poorly printed text. Using an 18th-century German passage about an orchid near Frankfurt, he showed the occurrence and trait records that could be extracted. He then showed a current LLM getting taxon and habitat right but hallucinating GBIF and Wikidata IDs, even after prompt improvements, so outputs must be verified against authoritative sources. He presented BIOfid, the DFG-funded specialised information service run by Senckenberg and Goethe University Frankfurt, with its annotated corpus, Unified Corpus Explorer, NLP pipelines, ontologies and plans for multimodal extraction into RDF/OWL knowledge graphs.
Why it matters. Centuries of occurrence and trait data exist only in print. Standard-based, verifiable extraction pipelines are needed before such data can safely feed knowledge graphs and AI.
In the room
- Collections data have depth but limited volume, and citizen science has volume but only recent years; literature combines both but is the hardest to access.
- Challenges: natural language, non-English texts (German, Latin, Hungarian), insider abbreviations, domain conventions such as indented keys, poor print quality.
- Mobilisation means converting analog, unstructured information into FAIR digital formats, not just digitisation.
- Example: an occurrence of Coeloglossum viride near Eschenheimer Tor, Frankfurt, May 1729, in a meadow, plus a leaf-width trait record, extracted from one sentence.
- An LLM extracted taxon and habitat correctly but hallucinated GBIF and Wikidata IDs, still wrong after a stricter prompt; GBIF and World Flora Online also disagree on the accepted name (Coeloglossum viride vs Dactylorhiza viridis).
- BIOfid aims: mobilise data from Central European literature, provide text-mining tools, digitise 20th-century literature; a collaboration of Senckenberg, Goethe University text technology and the Senckenberg Library, funded by DFG.
- Uses the generic standards UIMA, OWL and RDF/SPARQL and the domain standards Darwin Core and GBIF/Catalogue of Life backbones. Reusable components are DUUI, the Unified Corpus Explorer (UCE) and ontologies (e.g. Chilopoda anatomy, EnvO extensions and translations).
Notable moments
Transcript
Automatically generated captions can contain mistakes, especially in names and technical terms. Times are relative to the room recording.


