talk · Friday 25 September · SAL A

Using Large Language Models to enhance biodiversity knowledge extraction: From taxonomic treatments to graph-based knowledge systems

Ian Ondo · Large Language Models for Biodiversity Data Discovery, Integration, and Curation. Part A: AI- enabled knowledge creation and extraction

Recording time 1:37:16–1:50:58Open on Vimeo ↗

The short versionTaxonomic treatments hold underused habitat knowledge that LLMs and a provenance-preserving knowledge graph can structure and link to ecoregions, helping map data-poor species.

Overview

What this was about

Ian Ondo presented a framework that turns habitat information in taxonomic treatments into a structured, evidence-traceable knowledge base. Plazi TreatmentBank RDF becomes a source graph; an instruction-tuned LLM (Qwen 2.5) extracts explicit habitat statements into an evidence graph; the evidence is linked to environmental ontologies, WWF/One Earth terrestrial ecoregions and IUCN habitat classes. A trained habitat-environment compatibility model (HECC) put the correct ecoregion in its top three 75% of the time (top five 87%), and 65% zero-shot. A data-deficient Dalbergia species showed that the model's most compatible ecoregions matched its few known GBIF occurrences.

Why it matters. There is no comprehensive resource for species habitat preferences. Mining treatments with traceable evidence could support conservation assessments of data-deficient species.

Key ideas

In the room

  • Plazi's TreatmentBank has made over 1.2 million treatments from about 94,000 articles machine accessible, but their ecological content is not immediately reusable.
  • There are structured resources for taxonomy and occurrences, but nowhere comprehensive to ask what habitat a species prefers.
  • Framework: TreatmentBank RDF -> source graph (publication and treatment structure) -> AI habitat extraction -> evidence graph -> integration with environmental and spatial knowledge.
  • Existing standards and ontologies are reused wherever possible rather than new vocabularies.
  • Qwen 2.5 is prompted to extract only explicit habitat statements, not to infer from altitude, place names or coordinates, and to return nothing when there is none.
  • Around 40,000 treatment sections were processed. HECC, an embedding model, learns semantic compatibility between habitat text and ecoregion descriptions.
  • Ecoregion retrieval: correct in top 3 75%, top 5 87%; zero-shot ecoregions 65%.
Jump in

Notable moments

In their words

Transcript

Automatically generated captions can contain mistakes, especially in names and technical terms. Times are relative to the room recording.

Read the transcript ↓
Loading transcript…