Enhancing Biodiversity Georeferencing Using Large Language Models and Knowledge Graphs
Roselyn Gabud · Large Language Models for Biodiversity Data Discovery, Integration, and Curation. AI applications, integration, and validation
The short versionLLM preprocessing plus knowledge-base matching fills in locality hierarchies and proposes coordinates, with a review tool to keep humans in control.
What this was about
Roselyn Gabud presented a georeferencing pipeline for Manchester Museum herbarium site records from Africa exported from EMu. An LLM, backed by GeoNames and Wikidata, fills missing locality hierarchy fields from narrative site descriptions, with spelling correction and per-entity fallbacks. Geographic entities are then matched from the most to the least specific level and scored by name similarity and hierarchy. A human-in-the-loop tool shows suggestions on a map, lets annotators correct them, re-run georeferencing and export back to EMu. Preprocessing cut records with only one populated locality field from 134 to 10, and records with four or more fields rose from 13 to 142.
Why it matters. Hundreds of thousands of specimen records still await manual georeferencing; semi-automated methods that cope with sparse, ambiguous localities could clear that backlog.
In the room
- Collections such as Manchester Museum and Philippine biodiversity collections hold hundreds of thousands of records still georeferenced manually.
- Many records have only a continent and a narrative site name, so filtering by country or province is impossible; ambiguity comes from outdated names, spellings and overlapping names.
- Preprocessing: clean and extract location text, LLM interpretation of the full description, per-entity fallback via spelling correction, GeoNames, Wikidata and LLM, then hierarchy enrichment with confidence ratings.
- Georeferencing queries GeoNames, Wikidata or Wikipedia from nearest named place up to continent, scoring candidates on name similarity and hierarchy; no match leaves coordinates blank.
- Knowledge bases compared on 100 records (0.1-degree threshold): GeoNames good for administrative hierarchy but limited; Wikidata broader for natural features but noisy hierarchy; Wikipedia most comprehensive but noisy and slow.
- On 263 annotated records (50 train, 213 dev), completeness improved markedly after preprocessing.
Notable moments
Transcript
Automatically generated captions can contain mistakes, especially in names and technical terms. Times are relative to the room recording.


