A conference theme

ML training data and benchmarks

Building, sharing and evaluating training datasets and benchmarks for biodiversity machine learning, including use of museum images as training data.

13 talks & discussions
Across rooms and sessions

Explore the conversation.

talk · Monday 21 September · Aula

AI for Nature at Large Scale and High Resolution

AI can help fill biodiversity knowledge shortfalls only if biological knowledge is built into the models and the data infrastructure is FAIR for AI and returns value to primary data collectors, because 'there is no AI without data'.

Tanya Berger-Wolf ↗
talk · Tuesday 22 September · SAL A

AI-vision for integrated management of pest insects and their natural enemies

Moving precision spraying from static weeds to mobile pests and their natural enemies is blocked mainly by the difficulty of collecting representative image data.

Ritter Guimapi ↗
talk · Tuesday 22 September · SAL A

Are plant image datasets ready for AI-enabled community assessment? A systematic review of bias, representativeness and annotation quality

Current plant image datasets are not yet ready for AI-enabled ecological monitoring, especially for spatially explicit tasks and annotation quality.

Yuchen Zhao ↗
talk · Tuesday 22 September · SAL A

Cross-Modal Coordination for Biodiversity Monitoring at Scale: Lessons from SmartWilds

Coordinating multiple sensor modalities with AI at the far edge, informed by synchronised multimodal datasets like SmartWilds, can scale and adapt biodiversity monitoring.

Jenna Kline ↗
talk · Tuesday 22 September · SAL B

From satellite embeddings to Darwin Core: an open, ensemble-AI workflow for Prosopis juliflora mapping and reusable training-dataset generation

Satellite embeddings plus hyperspectral signatures and ground validation can map invasive Prosopis at scale and generate reusable, attributable training data.

Pranav Jha ↗
talk · Tuesday 22 September · SAL A

HerbAudit: Validating AI-Driven Herbarium Transcriptions

HerbAudit shows AI herbarium transcription reaching ~94% accuracy and provides a fair, field-aware way to benchmark it.

Dilara Ağacık ↗
talk · Tuesday 22 September · SAL B

KakraCards: An AI-Assisted Pipeline for Liberating Six Decades of Seabird Heritage Data

Multi-model consensus with human adjudication builds ground truth and picks the best LLM for transcribing standardised historical cards.

Kristjan Adojaan ↗
talk · Tuesday 22 September · SAL A

The WildLIVE Portal: Towards Scalable Wildlife Monitoring through the Integration of Citizen Science and Human-in-the-Loop AI Workflows

WildLIVE combines a model registry, provenance-tracked human verification and FAIR publication to move camera trap data from raw images to publishable data.

Rajapreethi Rajendran ↗
talk · Thursday 24 September · SAL A

Accelerating specimen identification through Virtual Collections while promoting collaboration

Virtual reference collections with DOIs let experts identify and curate dispersed digital specimens remotely instead of shipping material or travelling.

Melanie de Leeuw ↗
talk · Thursday 24 September · SAL B

AI Ready Standards with Croissant for Type Specimens Catalog Datasets

Combining LLM extraction, Darwin Core and Croissant metadata offers a route from historical specimen catalogues to FAIR, ML-ready datasets.

Sefika Efeoglu ↗
talk · Thursday 24 September · SAL B

Digital Twins as System-Optimization Tools: Provisioning Edge and Cloud Infrastructure for Biodiversity Monitoring at Scale

Digital twins should close the sim-to-real loop for the sensor systems too, letting teams plan where and how to deploy multimodal monitoring before going into the field.

Tanya Berger-Wolf ↗
talk · Thursday 24 September · SAL C

From 1.5 Billion Raw Queries to AI-Ready Biodiversity Data: Human-AI Collaborative Curation Pipelines in Pl@ntNet

Pl@ntNet's human–AI pipeline filters a vast, noisy stream into two GBIF datasets and training data, with separate branches for opted-in human-validated and automatic occurrence-only data.

Alexis Joly ↗
talk · Friday 25 September · SAL B

Beyond traditional identification keys: hybrid deep learning and Xper3 keys for insect identification

Deep-learning pre-filtering of an interactive key keeps the key's explainability while greatly shortening identification paths.

Agathe Puissant ↗