talk · Thursday 24 September · SAL A

Aligning An Evolving Global Data Model To Software With Global Ambitions: GBIF x TaxonWorks

Matthew Yoder · Solutions for Research Collection Management Systems Challenges

Recording time 6:38:30–6:50:46Open on Vimeo ↗

The short versionTaxonWorks and DwC-DP have converged conceptually, but exchange between rich models will always be lossy, so mappings should be recorded explicitly.

Overview

What this was about

Matt Yoder compared TaxonWorks' API-first data model with the new Darwin Core Data Package (DP). He found the conceptual overlap 'uncanny' after years of independent development. DP's many media, identifier, protocol and reference link tables collapse into single polymorphic annotation tables in TaxonWorks. TaxonWorks lacks a primary Survey concept and full nucleotide analysis, while DP lacks precise anatomy links and physical container hierarchies. He proposed recording mappings (broader/narrower) in SSSOM, merging them in RDF into a 'data mapping of all data mappings', and exposing /vocab endpoints. His observations: partial alignment and lossy exchange are the norm, both models have competing 'fuzzy buckets', and DP's role as an exchange format is unclear.

Why it matters. It is a concrete software–standard alignment exercise with ideas (SSSOM mappings, /vocab endpoints) that could help many CMSs map to DwC-DP. It also flags practical issues like duplicate extension buckets.

Key ideas

In the room

  • TaxonWorks serves taxonomists first and now about 10–11 large collections; it is API-first and set up for human and agentic contributions.
  • It already offers Darwin Core import/export, a built-in IPT-like GBIF endpoint, EML persistence and lossless DwC import; an Orthoptera project round-tripped over a million GBIF records.
  • Conceptual overlap with DwC-DP is remarkable; TaxonWorks models one image depicting many concepts via a single annotation pattern.
  • Gaps in TaxonWorks: Survey as a primary concept, full nucleotide analysis; gaps in DP: precise anatomy, physical containers, loans, stronger geo-reference primacy.
  • Mappings can be expressed as broader/narrower relations and persisted in SSSOM, merged in RDF.
  • Both models have two 'buckets' for fuzzy data (resource relationship vs assertions; observations), which will confuse users.
  • Needs identified: equipment and configuration, experiments, genomes, more strongly typed agents; comments used as definitions cause confusion.
Jump in

Notable moments

In their words

Transcript

Automatically generated captions can contain mistakes, especially in names and technical terms. Times are relative to the room recording.

Read the transcript ↓
Loading transcript…