Beyond traditional identification keys: hybrid deep learning and Xper3 keys for insect identification
Agathe Puissant · May the Data Be Structured: Linking Descriptions, Identification and AI
The short versionDeep-learning pre-filtering of an interactive key keeps the key's explainability while greatly shortening identification paths.
What this was about
Puissant presented a hybrid identification method. A supervised contrastive metric-learning model trained on photographs of MNHN collection specimens of French ladybirds (Coccinellidae) returns the nearest candidate taxa, which then pre-filter an Xper3 interactive key. The model's top-10 hit rate was about 90% on both unstandardised GBIF images and standardised professional photographs. Pre-filtering cut the mean number of key steps sharply, especially for difficult cases, while leaving the final decision to the user.
Why it matters. The approach turns digitised museum collections into training data for tools that non-experts, e.g. in large-scale monitoring, can use without handing the final decision to a black box.
In the room
- Collection specimens give precisely identified images covering broad morphological variation, but such datasets are small, have many classes and are imbalanced.
- Automated tools such as BirdNET and Pl@ntNet need large training sets and were described as usually less taxonomically precise than expert-built keys.
- Uses supervised contrastive learning (SupCon): a CNN embeds images so that same-class pairs are pulled together and different-class pairs pushed apart, which enables nearest-neighbour identification.
- Evaluation metric: hit rate, i.e. whether the correct taxon is among the 10 nearest neighbours. Hyperparameter search selected a deep network with long training and moderate augmentation.
- External validation on GBIF images (variable lighting and pose) and on a professional photo website both gave about a 90% hit rate.
- In an Xper3 interface the uploaded image filters the key. Performance drops for dark or complex backgrounds, since training images had backgrounds removed.
- The average path length fell from 4.2 steps to about one (ASR partly garbled); the longer the original path, the larger the reduction.
Notable moments
Transcript
Automatically generated captions can contain mistakes, especially in names and technical terms. Times are relative to the room recording.


