Panel discussion: From Mobilizing Data to AI-Ready Knowledge
Hilmar Lapp, Robert Guralnick, Tanya Berger-Wolf, Alexis Joly, Lars Vogt, Tim Alamenciak, Joe · From Mobilizing Data to AI-Ready Knowledge: Infrastructure for Multimodal Biodiversity Data
The short versionAI-readiness is purpose-dependent, and the hardest remaining problems are social: trust, attribution, governance and data sovereignty, where ontologies and knowledge graphs offer some control and explainability.
What this was about
The SYM25 speakers took questions on AI-ready biodiversity data, moderated by Hilmar Lapp. Panellists agreed that AI-readiness depends on purpose, like fitness for use. That is why FAIR4AI is distinguished from task-specific AI-readiness, though aggregation mixes purposes and shared benchmark snapshots remain valuable. On benefits to data providers, Tanya Berger-Wolf described dataset-side 'recipes' for agents, identifiers propagating from the smallest data unit through aggregations, and a CyVerse/NEON pilot where models register their use against operationalised permissions. Asked what is hardest now, panellists named reluctance to open data to tech-giant harvesting, pushback in Latin America and Africa over data sovereignty and digital sequence information, and data heterogeneity as the root of AI bias. A GBIF voice countered with strong demand for open data in the BID programme and Indigenous data work. A final question on ontologies drew views on ontology governance, explainability and knowledge graphs as the part of LLM systems that users can control.
Why it matters. The panel brought together the symposium's threads and set out open problems (value flowing back to data providers, sovereignty, heterogeneity, ontology governance) that infrastructures and standards bodies such as TDWG and GBIF will need to address.
In the room
- Warm-up: an audience member asked how to persuade data providers that their contributions will still be acknowledged when their data are used by AI agents, noting they had not seen a strong solution.
- AI-readiness depends on purpose: data are purpose-driven models of the world, and like data quality it is fitness for use, hence the FAIR4AI vs AI-ready distinction.
- Devil's advocate: aggregated datasets mix purposes, so each researcher must define fitness for their use; shared cleaned benchmark snapshots are still important.
- Bidirectionality: add 'recipes'/skills to datasets for agents; propagate identifiers from the smallest data unit through every aggregation so citations and value flow back; store data once and reference derivatives.
- The CyVerse pilot with NEON operationalises data use and licensing permissions per model and use case, with models registering use so attribution and valuation propagate back.
- Hardest challenges: people withholding data from tech-giant harvesting; open data framed as Western/colonial ('doctrine of digital discovery'); the DSI debate at CBD COP16 moving toward biodiversity digital information; >90% of biodiversity data from private, Indigenous and local community land.
- Counterpoint from GBIF: the BID call drew over 1,000 applications requesting €48 million, less than 5% fundable, all requiring open data. Indigenous communities are willing to share when done properly.
Notable moments
Transcript
Automatically generated captions can contain mistakes, especially in names and technical terms. Times are relative to the room recording.


