talk · Thursday 24 September · SAL A

Mapping the Relational Data Schema of the Diversity Workbench (DWB) to the Darwin Core Data Package Data Model to fit better to GBIF´s New Data Model

Tanja Melanie Weibulat · Solutions for Research Collection Management Systems Challenges

Recording time 7:31:35–7:38:45Open on Vimeo ↗

The short versionDwC-DP promises to carry the relational richness that previously pushed GFBio centres to ABCD, and a first DWB-to-DwC-DP prototype via the test IPT works.

Overview

What this was about

Tanja Weibulat explained that GFBio data centres chose ABCD because early Darwin Core couldn't express DWB's complex relations. Their pipeline exports and filters DWB modules (withholding restricted data, anonymising people, rounding coordinates) into PostgreSQL, then maps to ABCD via BioCASe for GBIF and NFDI4Biodiversity. Limitations were that GBIF didn't display multiple identifiers, persons, roles or relations, and data papers needed manual entry. They prototyped mapping the same PostgreSQL tables to DwC-DP in a test IPT: observations and physical specimens split into Occurrence and Material, event, identification and agents shared, and analysis and processing combined into Material Assertion. Next steps are a GBIF sandbox comparison, metadata mapping and extending to the GFBio portal.

Why it matters. It is a concrete ABCD-to-DwC-DP transition test that matters for the many German data centres still publishing ABCD.

Key ideas

In the room

  • DWB publishing used ABCD because Darwin Core was 'too slow' to capture relations.
  • Pipeline: DWB modules → filtered export (withhold, anonymise, round coordinates) → PostgreSQL per data package → BioCASe → ABCD XML harvested by GBIF and NFDI4Biodiversity.
  • GBIF did not display multiple global identifiers, persons with roles, or relations present in ABCD; data papers needed manual entry.
  • Only a test version of the IPT currently supports DwC-DP; source tables were accessed directly via SQL.
  • Mapping required splitting observation vs specimen data into Occurrence and Material and combining analysis/processing into Material Assertion.
  • Dataset metadata must still be entered manually in the IPT.
Jump in

Notable moments

In their words

Transcript

Automatically generated captions can contain mistakes, especially in names and technical terms. Times are relative to the room recording.

Read the transcript ↓
Loading transcript…