talk · Tuesday 22 September · SAL B

Open Digital Curation and Round-Tripping Biodiversity Data Enhancements Across Collection Infrastructures: Six Years and 50 Million Links in Bionomia

David Shorthouse · Responsible AI, Open Digital Curation, and Round-Tripping for Biodiversity Data

Recording time 5:29:52–5:44:06Open on Vimeo ↗

The short versionRound-tripping is much harder than expected, and data managers must fully populate identifiedBy so taxonomists' expertise is credited.

Overview

What this was about

In a pre-recorded talk, David Shorthouse questioned calls for taxonomists to train AI identification models without credit, and highlighted that 60% of preserved specimens in GBIF lack any identifiedBy value. He reviewed six years of Bionomia: services for individuals (alerts when specimens are used in research, Zenodo archives with DOIs) and for organisations (frictionless data packages of CSV files, badges), and the potential of the collector social network for attribution suggestions. He identified unstable occurrence identifiers as the main barrier to round-tripping, noted no CMS can record third-party attribution provenance, and discussed the Darwin Core Data Package's agent terms, predicting recordedByID and identifiedByID will become obsolete and calling for a TDWG task group on an agent role vocabulary.

Why it matters. A reflective, data-backed critique of how collections credit expertise and how AI and round-tripping can erode or build trust.

Key ideas

In the room

  • A 2025 paper trained a herbarium-image identifier to 90% accuracy and recommended taxonomists train models, which he framed as working oneself out of a career without credit.
  • About 60% of preserved specimens published to GBIF have no identifiedBy content, undermining efforts like the EU TETTRIs project.
  • Bionomia is in its sixth year; anyone with ORCID can link specimens to collectors and determiners; interface in seven languages; ~320 million records re-downloaded monthly from GBIF.
  • DwC-DP will let over-prescribed terms (recordedBy, identifiedBy) be expressed as agent-role-action relationships; he predicts recordedByID and identifiedByID will be deprecated, questions the 'preferred' prefix in agent name, and calls for a TDWG Attribution Interest Group task group to ratify agent roles (perhaps modelled on the OBO Contributor Role Ontology).
  • The co-collector social network can boost candidate matches and could support a lightweight attribution utility within a CMS.
  • User services: email alerts when one's specimens are used in research; Zenodo snapshots with DOIs for CVs. Organisation services: frictionless data package of CSVs (users, missing attributions, date conflicts) and HTML badges.
  • Main barrier: unstable occurrence identifiers from dataset deletions, splits and migrations; catalogue number is often the more stable key.
Jump in

Notable moments

In their words

Transcript

Automatically generated captions can contain mistakes, especially in names and technical terms. Times are relative to the room recording.

Read the transcript ↓
Loading transcript…