talk · Tuesday 22 September · ODIN

Resilient Attribution Infrastructures for Biodiversity Science: Archiving and Risk Mitigation in Bionomia

David Shorthouse · Building resilient data infrastructures for biodiversity science

Recording time 1:22:39–1:29:13Open on Vimeo ↗

The short versionSmall niche infrastructures can be resilient by planning for their own end: keep things simple and push versioned data and reusable code outward.

Overview

What this was about

David Shorthouse describes Bionomia, which links more than 50 million biodiversity records to people via ORCID and Wikidata identifiers through automation and volunteer 'scribes', and the risks it faces: code and server upkeep, institutional oversight questions with his employer, and AI crawlers that took the site down dozens of times until Cloudflare was introduced. Resilience comes from simple design, standalone libraries (a Ruby gem ported to Go for parsing human names), and monthly archiving to Zenodo of a CSV graph and frictionless data packages for roughly 9,000 GBIF datasets, which institutions such as the Natural History Museum, University of Oslo reuse. Small projects can be resilient if they plan for decommissioning from the outset.

Why it matters. Bionomia is a widely used single-maintainer service; its approach shows how to protect volunteers' contributions and downstream users against a project's eventual loss.

Key ideas

In the room

  • Bionomia maintains over 50 million attribution links; interface translated into seven languages via Crowdin.
  • Runs on three DigitalOcean virtual machines with a Redis job queue.
  • Misbehaving bots and AI crawlers took the website down dozens of times; Cloudflare now gatekeeps, but its long-term sustainability is a worry.
  • Core functions are encapsulated in standalone libraries, e.g. a Ruby gem for parsing human names now ported to Go.
  • Monthly: a CSV graph structure pushed to Zenodo, and a frictionless data package for each of about 9,000 GBIF-publishing datasets, including logically incongruent records (e.g. collected before a collector's birth).
  • Data round-trips: e.g. NHM University of Oslo staff reuse the packages and publish ORCID/Wikidata identifiers to GBIF.
  • Individual researchers can archive their attributed records with DataCite DOIs via Zenodo integration.
Jump in

Notable moments

In their words

Transcript

Automatically generated captions can contain mistakes, especially in names and technical terms. Times are relative to the room recording.

Read the transcript ↓
Loading transcript…