talk · Thursday 24 September · SAL C

Local scientific name resolution using custom data sources

Dmitry Mozzherin · Digital Tools for Data Discovery, Resolution and Exchange

Recording time 8:49:02–9:00:07Open on Vimeo ↗

The short versionGNverifier can now be run locally over any custom or private name datasets, using the standardised SFGA format and GNdb, with the same API as the central service.

Overview

What this was about

Dmitry Mozzherin presented GNverifier, a scientific name reconciliation tool with fuzzy, partial and authorship matching, over 80 million pre-indexed records from more than 100 datasets, and throughput above 2,000 names per second. It can now run locally as well as at verifier.globalnames.org. Local use suits proprietary, institutional or specialised datasets, offline fieldwork, frequent updates and faster, slimmer databases. The new GNdb tool, part of the Species File Group's SFWork approach, ingests Species File Group Archive (SFGA) files. SFGA can be used directly without database conversion, and the sf and harvester programs produce it from Darwin Core archives, ColDP, CSV/TSV or plain name lists. This replaces error-prone ad-hoc parsing and ingests Catalogue of Life in minutes rather than hours.

Why it matters. Local, reproducible name verification against private or project-specific checklists supports data sovereignty and field use, a need that centralised services cannot meet.

Key ideas

In the room

  • GNverifier does partial, fuzzy and authorship matching over >80 million records from >100 datasets at >2,000 names per second.
  • It runs locally or via the central service at verifier.globalnames.org with the same workflow.
  • Reasons for local use: datasets missing from the central service (proprietary, institutional, specialised), offline/field use, updating datasets at will, and including only needed datasets for speed.
  • Previous ingestion of diverse, non-standard formats was painful (encoding, malformed metadata XML, delimiter and field-count errors).
  • GNdb, part of SFWork, ingests SFGA (Species File Group Archive) files, which are usable immediately without conversion into a database.
  • Pipeline: sf / harvester convert DwC-A, ColDP, CSV/TSV and text to SFGA → gndb populate into Postgres (custom datasets need YAML metadata and IDs ≥1000) → gndb optimize (denormalised view) → gnames service → gnverifier web/CLI/API.
  • Catalogue of Life can now be ingested in minutes instead of hours.
Jump in

Notable moments

In their words

Transcript

Automatically generated captions can contain mistakes, especially in names and technical terms. Times are relative to the room recording.

Read the transcript ↓
Loading transcript…