Resilient Attribution Infrastructures for Biodiversity Science: Archiving and Risk Mitigation in Bionomia
David Shorthouse · Building resilient data infrastructures for biodiversity science
The short versionSmall niche infrastructures can be resilient by planning for their own end: keep things simple and push versioned data and reusable code outward.
What this was about
David Shorthouse describes Bionomia, which links more than 50 million biodiversity records to people via ORCID and Wikidata identifiers through automation and volunteer 'scribes', and the risks it faces: code and server upkeep, institutional oversight questions with his employer, and AI crawlers that took the site down dozens of times until Cloudflare was introduced. Resilience comes from simple design, standalone libraries (a Ruby gem ported to Go for parsing human names), and monthly archiving to Zenodo of a CSV graph and frictionless data packages for roughly 9,000 GBIF datasets, which institutions such as the Natural History Museum, University of Oslo reuse. Small projects can be resilient if they plan for decommissioning from the outset.
Why it matters. Bionomia is a widely used single-maintainer service; its approach shows how to protect volunteers' contributions and downstream users against a project's eventual loss.
In the room
- Bionomia maintains over 50 million attribution links; interface translated into seven languages via Crowdin.
- Runs on three DigitalOcean virtual machines with a Redis job queue.
- Misbehaving bots and AI crawlers took the website down dozens of times; Cloudflare now gatekeeps, but its long-term sustainability is a worry.
- Core functions are encapsulated in standalone libraries, e.g. a Ruby gem for parsing human names now ported to Go.
- Monthly: a CSV graph structure pushed to Zenodo, and a frictionless data package for each of about 9,000 GBIF-publishing datasets, including logically incongruent records (e.g. collected before a collector's birth).
- Data round-trips: e.g. NHM University of Oslo staff reuse the packages and publish ORCID/Wikidata identifiers to GBIF.
- Individual researchers can archive their attributed records with DataCite DOIs via Zenodo integration.
Notable moments
Transcript
Automatically generated captions can contain mistakes, especially in names and technical terms. Times are relative to the room recording.


