talk · Thursday 24 September · SAL C

From Legacy Data to FAIR Workflows: Standardising Biodiversity Collections Using Specify

Siyabonga Zamisa · Meeting Biodiversity Data Standards Using Specify Software

Recording time 1:17:48–1:30:23Open on Vimeo ↗

The short versionStandardised workflows, not just a new database, turned 11 differently managed collections into GBIF-ready data.

Overview

What this was about

Siyabonga Zamisa described how the KwaZulu-Natal Museum moved its 11 natural history collections (over 430,000 records) from Access, Excel, catalogue cards and other systems into Specify by 2019. With support from the Natural Science Collections Facility, it then developed shared workflows. He walked through a Darwin Core-based template spreadsheet used by all collections, quality checks in OpenRefine before WorkBench upload, routine specimen-versus-database QA, and a publishing pipeline through the IPT to GBIF. Before this the museum had no collections in GBIF; seven of eleven are now published.

Why it matters. This is a concrete African museum example of procedure and standards, not technology alone, driving data mobilisation, with measurable gains in publishing, cataloguing speed and loan processing.

Key ideas

In the room

  • 11 collections with over 430,000 records, the largest being Mollusca and Insecta. Previous systems included Access, Excel and Rapid, and all collections have been in Specify since 2019.
  • The problem was inconsistent practice: each collection had its own habits for capturing, tracking and sharing data, which slowed loans and data sharing.
  • A single template spreadsheet based on Darwin Core terms, with field explanations, is used across all 11 collections to ease WorkBench mapping.
  • Workflow: capture in template → data technician QC in OpenRefine → Specify WorkBench mapping and validation → upload.
  • Daily QA: randomly selected specimens are checked against their Specify records and both are corrected.
  • Publishing: Specify query export → OpenRefine QC → removal of sensitive species → IPT metadata and Darwin Core mapping → GBIF.
  • Results: from no collections in GBIF to seven of eleven, faster cataloguing and much shorter loan processing.
Jump in

Notable moments

In their words

Transcript

Automatically generated captions can contain mistakes, especially in names and technical terms. Times are relative to the room recording.

Read the transcript ↓
Loading transcript…