talk · Tuesday 22 September · ODIN

How Darwin Core enables harmonizing Biodiversity Data via the OSCA Consortium

Cristian-Dan Bara · Data Quality: From Standard to Practice

Recording time 3:45:13–3:58:54Open on Vimeo ↗

The short versionA shared Darwin Core-based cloud pipeline lets 15 Austrian institutions of very different IT capacity publish harmonised specimen data.

Overview

What this was about

Cristian-Dan Bara presents OSCA (Open Scientific Collections Austria), a collaboration of 15 Austrian institutions digitizing preserved specimens into a shared cloud infrastructure built around a central Postgres cache aligned to Darwin Core and approved extensions. A streamlined pipeline with automated checks for Darwin Core validity and essential values gives scientists feedback before publication, keeping a human in the loop. Two public portals (v4 aggregating searches across international portals, v5 searching consortium data first with deep links outward) are being compared, and records are categorised by MIDS levels, targeting at least MIDS level 1.

Why it matters. It shows how a national consortium can use Darwin Core as a common language to pool capacity and publish comparable collection data.

Key ideas

In the room

  • OSCA digitizes preserved specimens across 15 museums, universities and scientific institutions.
  • Darwin Core is the shared language, enabling "micro infrastructures" where larger institutions support smaller ones.
  • Central Postgres table in the cloud follows Darwin Core and approved extensions.
  • Automated processing flags Darwin Core validity and missing essential values; scientists correct before publication.
  • Three pillars: search and discover, direct Darwin Core export (Excel), and a feedback loop on user needs.
  • Portal v4 compiles results from international portals; v5 searches consortium data first (harvested into DiSSCo) with deep links.
  • MIDS classes used to categorise records; MIDS level 1 minimum, level 2 (with media) desired for public value.
Jump in

Notable moments

In their words

Transcript

Automatically generated captions can contain mistakes, especially in names and technical terms. Times are relative to the room recording.

Read the transcript ↓
Loading transcript…