discussion · Friday 25 September · SAL B

Discussion: designing institutional biodiversity knowledge and data science centres

Unidentified session chair, Lars Vogt, Tarek Al Mustafa, Matthew Yoder, Deb Paul, Patricia Mergen, Tim, Nikki, Sharon, Noah · Designing Institutional Knowledge Data Science Centres for Biodiversity

Recording time 4:04:57–4:25:14Open on Vimeo ↗

The short versionInteroperable institutional knowledge spaces depend less on technology than on people, shared storage and semantics, incentives and governance.

Overview

What this was about

After summarising the session's talks, the chair asked four questions. What belongs in an institutional centre and what in shared federated infrastructure? What is the biggest obstacle? What should change because of AI? What must institutions agree on to interoperate in 5-7 years? Answers ranged from data-owner-controlled federation, with institutions stewarding shared semantic artefacts, to calls for radical sharing of compute, tools and data against corporate interests. Others raised access inequalities, Indigenous data governance, researcher incentives and journal enforcement, tools that hide semantics, scouting community innovation, embracing messiness, US-Europe differences in AI policy, investing in people, and archival format portability.

Why it matters. The discussion gives a cross-institutional view of priorities and tensions (openness versus governance, European versus US AI policy, incentives) for anyone planning biodiversity knowledge centres.

Key ideas

In the room

  • Vogt: hosting depends on data owners. Some want full control via federated nodes that define allowed queries; others need institutes to host for them. Semantic artefacts (ontologies, schemata) and core services should sit with institutions for long-term sustainability. The chair compared this to GBIF nodes.
  • An audience member noted that skills (not only data) should be shared, and that recognition of data and software skills is patchy across countries.
  • Tim argued for radical sharing of compute, tools and resources, since large AI companies and big publishers benefit from secrecy and concentration.
  • Deb Paul noted access inequalities and her institution's experience that many researchers had never tried AI. She proposed Carpentries-style 'day three' support and alliances to handle staff turnover.
  • Sharon (Field Museum) urged bringing Indigenous data governance and traditional knowledge into 'make everything open'. A Brazilian participant disputed a claim about smartphone access in Brazil.
  • Obstacles: Tarek said researchers lack incentives and skills to semantify data, so institutions and computer scientists must build supporting workflows. Another participant said researchers know how but take shortcuts, so journals should require standardised data. Matt said to build tools that hide the semantic layer, and that federated storage is the biggest obstacle. A further point was that reviewers should flag messy data.
  • On AI: Matt called for small teams trained to spot community innovations and bring them into standards. Tim suggested embracing authentic messiness (citing Journal of Open Source Software practice). Noah flagged a gap between European institutions restricting big-tech AI and widespread US academic use.
Jump in

Notable moments

In their words

Transcript

Automatically generated captions can contain mistakes, especially in names and technical terms. Times are relative to the room recording.

Read the transcript ↓
Loading transcript…