Curation through citation: using AI and a knowledge graph to curate DNA barcodes
Roderic Page · Large Language Models for Biodiversity Data Discovery, Integration, and Curation. AI applications, integration, and validation
The short versionLinking barcodes, specimens and literature in a knowledge graph could fix many BOLD records, if we can reliably extract the connecting codes from papers.
What this was about
Roderic Page asked what concrete value knowledge graphs deliver, and proposed curation of DNA barcode data as an answer. Using his BOLD View site, he showed BOLD records with bad or missing coordinates, vague or informal names, and names hidden behind GenBank links. Some can be fixed by following links from the barcode to GenBank, the publishing paper's DOI and a holotype, which is more than internal-consistency data cleaning. The barriers are extracting specimen and sequence codes from literature (which Plazi does not do well) and linking barcodes to papers. He pointed to emerging large triple stores (Wikidata, OpenStreetMap, GBIF, BHL, Bionomia; 86 billion triples) as signs that such a graph is becoming real.
Why it matters. Barcode data train widely used identification models while many records lack species names or correct coordinates. Citation-based curation would improve data that downstream AI relies on.
In the room
- Knowledge graphs need to show 'concrete information you do not already possess' (the Silicon Valley quote).
- CurateGPT suggested combining knowledge graphs and chatbots for curation.
- BOLD holds ~25 million barcodes; a well-known 5-million-barcode training set is mostly identified only to order or family.
- BOLD View is his site for exploring barcode data and noting odd records.
- Examples: a Tasmanian-endemic crustacean plotted at Australia's centroid; a record labelled only 'Caudata' that GenBank calls a snake from China; an informal name resolved via GenBank -> DOI -> paper -> holotype.
- Data cleaning flags inconsistencies; fixing needs external evidence such as the literature.
- Plazi extracts coordinates well but not specimen and sequence codes; a simple ChatGPT prompt gets most of them.
Notable moments
Transcript
Automatically generated captions can contain mistakes, especially in names and technical terms. Times are relative to the room recording.


