Are plant image datasets ready for AI-enabled community assessment? A systematic review of bias, representativeness and annotation quality
Yuchen Zhao · AI for Biodiversity Data
The short versionCurrent plant image datasets are not yet ready for AI-enabled ecological monitoring, especially for spatially explicit tasks and annotation quality.
What this was about
Yuchen Zhao (Newcastle University) argued that ecological AI is constrained less by model architecture than by the quantity and quality of training data. She presented a four-dimension framework, inspired by clinical reporting frameworks and FAIR, for assessing plant image dataset readiness: training-relevant characteristics, image acquisition characteristics, annotation protocol, and accessibility. A systematic review of in-situ, non-crop plant object detection and segmentation datasets screened 360 unique studies and found only 40 eligible datasets, with strong geographic bias toward Europe and Asia, a dominance of agricultural applications (19 of 40 focused on wheat), neglect of background plants, and generally unsatisfactory annotation protocols.
Why it matters. Spatially explicit tasks (locating plants and estimating cover) are needed for ecological inference, and the review offers a framework for building and reporting datasets fit for them.
In the room
- Imaging tasks: classification (present?), object detection (where?), segmentation (where exactly, and what cover?).
- Framework dimensions: training-relevant characteristics; acquisition (growth stage, season, angle, location, protocol, weather); annotation protocol and QC; accessibility.
- 508 records, 360 unique studies screened, 40 eligible datasets.
- Object detection datasets are increasing; segmentation less so; ground-level and UAV imagery dominate with growing use of consumer devices.
- Datasets are uneven in size and concentrated in Europe and Asia, neglecting biodiversity hotspots.
- Wheat (19/40) and tree crowns (8) dominate; only 3 on wildflowers.
Notable moments
Transcript
Automatically generated captions can contain mistakes, especially in names and technical terms. Times are relative to the room recording.


