talk · Tuesday 22 September · SAL A

Scalable Edge AI for Real-Time Biodiversity Monitoring: tracking invasive plant species in roadside imagery

Giulio Martellucci · AI for Biodiversity Data

Recording time 28:33–41:47Open on Vimeo ↗

The short versionDistilling a tiling-based plant identifier into a single ConvNeXt model makes high-resolution roadside invasive-species monitoring fast, cheap and more accurate.

Overview

What this was about

Giulio Martellucci presented work from a Biodiversa+ project in which vehicle-mounted cameras captured high-resolution roadside imagery across several countries (about 2,200 km of roads, 60 target alien plant species). Applying the Pl@ntNet classifier with multi-scale tiling works but takes around 10 seconds per image, so the team used knowledge distillation from the tiling 'teacher' to a global multi-label model: first the DINOv2 vision transformer, then a ConvNeXt (DINOv3-pretrained) model that scales linearly with resolution. The ConvNeXt model was both far cheaper (about 12 hours and 2.6 kWh versus 18 days and 138 kWh for the monitored roads) and more accurate on an expert-annotated test set.

Why it matters. Transport corridors spread invasive plants; efficient models that could run in real time and be released open source make continental-scale early detection feasible.

Key ideas

In the room

  • Invasive plants are among the top five drivers of biodiversity loss; roads and railways accelerate seed dispersal.
  • Multi-scale tiling with the Pl@ntNet API (~70,000 species) gives presence/absence but is slow and costly.
  • Knowledge distillation trains a student to mimic the tiling teacher's predictions on whole high-resolution images.
  • Vision transformers scale quadratically with resolution; ConvNeXt scales linearly.
  • ConvNeXt processes about 14 images per second and uses dramatically fewer GFLOPs than tiling.
  • ConvNeXt achieved the best mAP and ROC metrics on a test set of 36 monitored species annotated by experts; tiling ranks true occurrences but produces too many false alarms.
  • Grad-CAM and dense-feature analyses show the model localises species even with occlusion and overlap.
Jump in

Notable moments

In their words

Transcript

Automatically generated captions can contain mistakes, especially in names and technical terms. Times are relative to the room recording.

Read the transcript ↓
Loading transcript…