A metadata-first FAIR4AI checklist and automated evaluation agent for biodiversity datasets
Tanya Berger-Wolf · From Mobilizing Data to AI-Ready Knowledge: Infrastructure for Multimodal Biodiversity Data
The short versionFAIR4AI is mainly a metadata gap, and a community checklist plus an automated agent can measure and help close it while tracking provenance and governance back to data sources.
What this was about
Tanya Berger-Wolf presented the NSF-funded FAIR for AI Bio project (FAIROS programme), which has built AI-ready data infrastructure for ecology through a series of community workshops since August 2025. The project distinguishes FAIR, FAIR4AI (machine-findable, accessible, interoperable and reusable, so AI-ready datasets can be prepared with minimal human intervention) and AI-ready (application-specific) data, within a question–data–model round-tripping framework. The team built a checklist that grew from four categories (structural, scientific, provenance, governance) to 96 items, and an LLM agent that scores datasets, gives item-level recommendations and compares FAIR and FAIR4AI scores. Run on NEON datasets, the scores largely track each other but most datasets score very low on CARE.
Why it matters. As biodiversity data are fed into AI at scale, provenance, attribution and governance break easily. A measurable intermediate standard gives data providers and aggregators a concrete target.
In the room
- Rapidly growing sensor types and modalities are spread across many holders; integration should be decentralised rather than 'one database to rule them all'.
- 'There is no AI without data'; nature knows no borders but policies and licensing do.
- The FAIR for AI Bio project (NSF FAIROS) runs expanding workshops with GBIF, Movebank, Smithsonian, NatureServe, iNaturalist and others. The next workshop is February 2027 in California, with funding to bring participants.
- Data sovereignty and value round-tripping emerged as primary use cases at a conservation technology conference in Lima.
- Three tiers: FAIR; FAIR4AI (machine-actionable metadata enabling preparation of AI-ready data with minimal human intervention); AI-ready (application-specific).
- The checklist (developed starting with ESIIL) has four top-level categories (structural, scientific, provenance, governance) expanded to 96 items, and is open to community comment.
- An agent evaluates datasets against the checklist, giving scores, item-level recommendations, strengths and gaps, side by side with traditional FAIR.
Notable moments
Transcript
Automatically generated captions can contain mistakes, especially in names and technical terms. Times are relative to the room recording.


