From 1.5 Billion Raw Queries to AI-Ready Biodiversity Data: Human-AI Collaborative Curation Pipelines in Pl@ntNet
Alexis Joly · From Mobilizing Data to AI-Ready Knowledge: Infrastructure for Multimodal Biodiversity Data
The short versionPl@ntNet's human–AI pipeline filters a vast, noisy stream into two GBIF datasets and training data, with separate branches for opted-in human-validated and automatic occurrence-only data.
What this was about
Alexis Joly described the curation pipeline that turns Pl@ntNet's 1.6 billion raw, highly opportunistic image queries (from about 20 million users a year, including house plants, tattoos, people and sensitive content) into AI training data and GBIF occurrences. Every image passes a main model trained on about 25 million human-validated images, which includes 'reject' classes. He called that reject classifier the hardest part and an editorial choice. Opt-in public observations enter human curation with weighted majority voting, where each user's weight reflects expertise estimated by an EM algorithm. Other observations contribute only automatically filtered occurrences above a strict confidence threshold (<0.01% error) and a wildness classifier to GBIF. He warned about feedback loops in which users follow AI errors.
Why it matters. Opportunistic AI apps are among the largest sources of biodiversity observations. How they filter, validate, licence and publish data determines the quality of downstream models and species distribution analyses.
In the room
- Pl@ntNet is 15 years old, has about 20 million users a year, 1.6 billion raw observations and about 20,000 plants per hour.
- The raw stream includes cultivated and potted plants, plants on clothing and tattoos, humans (e.g. classroom sessions), non-living objects, nudity and drug-related images.
- The main model, trained on ~25 million human-validated images, outputs species, genus, family, diseases and reject classes. The reject classifier is the hardest and involves editorial choices about user intent.
- Branch 1: observations users opt in to share publicly (CC BY-SA, named) are curated by weighted majority voting, with user expertise estimated by an EM algorithm based on the number of species a user can recognise.
- Validated data retrain the model regularly, and old queries are reprocessed.
- Branch 2: non-shared observations contribute only occurrences (no images) above a per-model threshold with <0.01% error; about 150 million remain after this filter.
- A wildness classifier (trained on anthropised context and the app's cultivated tag) and WCVP level-2 country flora filtering remove cultivated plants before GBIF publication; two GBIF datasets result.
Notable moments
Transcript
Automatically generated captions can contain mistakes, especially in names and technical terms. Times are relative to the room recording.


