talk · Tuesday 22 September · ODIN

Rebuilding core functionalities of the GGBN system: taxonomic alignment with international research infrastructures

Walter Berendsohn · Semantic Interoperability and Standards Infrastructure

Recording time 5:42:25–5:56:23Open on Vimeo ↗

The short versionGGBN is rebuilding transparent automated name parsing and matching on GBIF parsing and ChecklistBank/Catalogue of Life XR to improve sample discovery.

Overview

What this was about

Walter Berendsohn describes replacing the outdated taxonomic component of the Global Genome Biodiversity Network (GGBN) portal, built around 2010. Incoming names from biobank datasets are first parsed (GBIF's parser chosen over Catalogue of Life's after testing), with pre-processing such as removing 'sp.' and flagging 'cf.', then matched via ChecklistBank name matching against the Catalogue of Life Extended Release to obtain synonyms and a standard higher taxonomy. Outcomes are qualified by match types and issues, with provider feedback for ambiguous cases, and the website documents how an accepted name was reached. He pleads for a universal homonym list.

Why it matters. Reliable taxonomic alignment across all kingdoms is essential to find biobank samples, do gap analysis and link to other infrastructures.

Key ideas

In the room

  • GGBN is a community instrument for DNA-research samples, not only a data portal; a white paper on GGBN was published this month.
  • Names should be processed automatically so data go online fast, while feedback to providers continues transparently.
  • Goals: better searching by any related name and by higher taxa, and future gap analysis.
  • GBIF and Catalogue of Life parsers share a core but give different results; GBIF's was chosen.
  • Parser attributes (e.g. virus, cultivar, Candidatus) indicate the nomenclatural code.
  • ChecklistBank name matching against Catalogue of Life Extended Release provides synonyms, broad coverage and standard higher taxa.
  • Pre-processing (removing sp., flagging cf.) significantly improves results.
Jump in

Notable moments

In their words

Transcript

Automatically generated captions can contain mistakes, especially in names and technical terms. Times are relative to the room recording.

Read the transcript ↓
Loading transcript…