talk · Friday 25 September · SAL A

Enhancing Biodiversity Georeferencing Using Large Language Models and Knowledge Graphs

Roselyn Gabud · Large Language Models for Biodiversity Data Discovery, Integration, and Curation. AI applications, integration, and validation

Recording time 4:02:43–4:14:00Open on Vimeo ↗

The short versionLLM preprocessing plus knowledge-base matching fills in locality hierarchies and proposes coordinates, with a review tool to keep humans in control.

Overview

What this was about

Roselyn Gabud presented a georeferencing pipeline for Manchester Museum herbarium site records from Africa exported from EMu. An LLM, backed by GeoNames and Wikidata, fills missing locality hierarchy fields from narrative site descriptions, with spelling correction and per-entity fallbacks. Geographic entities are then matched from the most to the least specific level and scored by name similarity and hierarchy. A human-in-the-loop tool shows suggestions on a map, lets annotators correct them, re-run georeferencing and export back to EMu. Preprocessing cut records with only one populated locality field from 134 to 10, and records with four or more fields rose from 13 to 142.

Why it matters. Hundreds of thousands of specimen records still await manual georeferencing; semi-automated methods that cope with sparse, ambiguous localities could clear that backlog.

Key ideas

In the room

  • Collections such as Manchester Museum and Philippine biodiversity collections hold hundreds of thousands of records still georeferenced manually.
  • Many records have only a continent and a narrative site name, so filtering by country or province is impossible; ambiguity comes from outdated names, spellings and overlapping names.
  • Preprocessing: clean and extract location text, LLM interpretation of the full description, per-entity fallback via spelling correction, GeoNames, Wikidata and LLM, then hierarchy enrichment with confidence ratings.
  • Georeferencing queries GeoNames, Wikidata or Wikipedia from nearest named place up to continent, scoring candidates on name similarity and hierarchy; no match leaves coordinates blank.
  • Knowledge bases compared on 100 records (0.1-degree threshold): GeoNames good for administrative hierarchy but limited; Wikidata broader for natural features but noisy hierarchy; Wikipedia most comprehensive but noisy and slow.
  • On 263 annotated records (50 train, 213 dev), completeness improved markedly after preprocessing.
Jump in

Notable moments

In their words

Transcript

Automatically generated captions can contain mistakes, especially in names and technical terms. Times are relative to the room recording.

Read the transcript ↓
Loading transcript…