Toward Quantitative AI-Readiness Metrics for Biodiversity Data Infrastructures
Yasin Bakış · AI-Readiness Metrics and Metadata for Biodiversity Data
The short versionAI readiness should be a task-specific, measurable profile rather than a single score, and current biodiversity standards lack most of the terms needed to describe it.
What this was about
Yasin Bakış defined AI readiness as the degree to which a dataset and its workflows can support a defined AI task reliably, reproducibly and interoperably, and argued it should be task-specific, multi-dimensional, testable and expressed as a profile rather than a single global score. Drawing on work preparing museum fish images from GBIF, iDigBio and museums for machine learning, he described moving from manual to automated metadata capture and feeding task performance (e.g. from a trait-extraction landmarking pipeline) back into the profile. Finding that about 90-95% of their roughly 200 needed terms were absent from Darwin Core, Audiovisual Core, EXIF and PROV, the team built the Fish-AIR platform and vocabulary and a TDWG task group under Audiovisual Core.
Why it matters. It makes AI readiness operational and exposes a concrete vocabulary gap in TDWG standards for image-based machine learning, with a route to address it via a task group.
In the room
- Definition: the degree to which a dataset and its associated workflows can support a defined AI task reliably, reproducibly and interoperably.
- AI readiness is task-specific, multi-dimensional, testable and profile-based; there should be a readiness profile, not a global score.
- Profile dimensions include metadata completeness, semantic alignment, provenance, image (data) suitability, reproducibility, interoperability and task performance.
- Metadata completeness is not just the share of filled columns but whether the fields ML practitioners use for filtering and cleaning are present.
- Manual metadata capture was too slow and expensive, so automated metadata capture was adopted.
- An AI-ready dataset is expected to perform better but does not guarantee good results; task performance scores from a trait-extraction/landmarking pipeline update the profile.
- Of roughly 200 terms developed, about 90-95% were new, not found in Darwin Core, Audiovisual Core, EXIF, PROV or other standards.
Notable moments
Transcript
Automatically generated captions can contain mistakes, especially in names and technical terms. Times are relative to the room recording.


