Geonomia: Metadata and Community for Georeferencing
Nicky Nicolson · Responsible AI, Open Digital Curation, and Round-Tripping for Biodiversity Data
The short versionClustering specimens into collecting trips makes georeferencing and AI enrichment more efficient and should be done as a community effort.
What this was about
Nicky Nicolson presented Geonomia, a Kew project to make specimens' textual localities computable. Only 13% of Kew's 6.3 million digitised specimens appear on GBIF maps, yet analysis of download predicates behind citations showed two-thirds of users filter spatially, discarding most of Kew's data. Using density-based clustering of collector number and date (collectors move forward in time and space), the team groups specimens into collecting trips across all herbaria's GBIF data, finds outliers, and passes trip-level batches to an LLM to summarise itineraries, habitats and regions, publishing results with Datasette. She argued collecting trips are a useful entity between GBIF's related records and Bionomia's collector careers, and that georeferencing should be shared community work.
Why it matters. Tackles the large share of specimen data invisible to spatial queries and offers a reusable approach for batching records meaningfully before applying AI.
In the room
- Kew has digitised 6.3 million specimens, but only 13% have coordinates on GBIF; many predate GPS and only have place descriptions.
- Analysis of GBIF download predicates for citations of Kew data showed two-thirds used spatial predicates, discarding most Kew data; the same analysis is possible for any GBIF dataset.
- Collector number plus date traces a collector's path; density-based clustering finds adjacent collecting events, groups variant locality strings and reveals transcription outliers.
- Clusters give context for human georeferencing and tighter context batches for AI.
- Used GBIF SQL downloads (vascular plants, species level, one country) to pre-process on GBIF's side, including screening recordedBy values like 'collector unknown'.
- Processing ran on a shared HPC facility at the Crop Diversity Institute; LLMs summarised habitats, itineraries (linear traverse vs hub-and-spoke), place names and regions per trip.
- Results published as a SQLite database with Datasette; example: a 1964 trip of almost 2,000 specimens with only about a third georeferenced, drawn from many institutions' datasets.
Notable moments
Transcript
Automatically generated captions can contain mistakes, especially in names and technical terms. Times are relative to the room recording.


