talk · Friday 25 September · SAL A

Curation through citation: using AI and a knowledge graph to curate DNA barcodes

Roderic Page · Large Language Models for Biodiversity Data Discovery, Integration, and Curation. AI applications, integration, and validation

Recording time 3:24:59–3:35:04Open on Vimeo ↗

The short versionLinking barcodes, specimens and literature in a knowledge graph could fix many BOLD records, if we can reliably extract the connecting codes from papers.

Overview

What this was about

Roderic Page asked what concrete value knowledge graphs deliver, and proposed curation of DNA barcode data as an answer. Using his BOLD View site, he showed BOLD records with bad or missing coordinates, vague or informal names, and names hidden behind GenBank links. Some can be fixed by following links from the barcode to GenBank, the publishing paper's DOI and a holotype, which is more than internal-consistency data cleaning. The barriers are extracting specimen and sequence codes from literature (which Plazi does not do well) and linking barcodes to papers. He pointed to emerging large triple stores (Wikidata, OpenStreetMap, GBIF, BHL, Bionomia; 86 billion triples) as signs that such a graph is becoming real.

Why it matters. Barcode data train widely used identification models while many records lack species names or correct coordinates. Citation-based curation would improve data that downstream AI relies on.

Key ideas

In the room

  • Knowledge graphs need to show 'concrete information you do not already possess' (the Silicon Valley quote).
  • CurateGPT suggested combining knowledge graphs and chatbots for curation.
  • BOLD holds ~25 million barcodes; a well-known 5-million-barcode training set is mostly identified only to order or family.
  • BOLD View is his site for exploring barcode data and noting odd records.
  • Examples: a Tasmanian-endemic crustacean plotted at Australia's centroid; a record labelled only 'Caudata' that GenBank calls a snake from China; an informal name resolved via GenBank -> DOI -> paper -> holotype.
  • Data cleaning flags inconsistencies; fixing needs external evidence such as the literature.
  • Plazi extracts coordinates well but not specimen and sequence codes; a simple ChatGPT prompt gets most of them.
Jump in

Notable moments

In their words

Transcript

Automatically generated captions can contain mistakes, especially in names and technical terms. Times are relative to the room recording.

Read the transcript ↓
Loading transcript…