talk · Tuesday 22 September · SAL A

Towards Scalable Trait Measurements from Digitized Specimens

Kenzo Milleville · AI for Biodiversity Data

Recording time 3:34:17–3:47:57Open on Vimeo ↗

The short versionMatching rulers to an annotated reference set lets specimen images be calibrated to real-world units with under 1% error in 96% of cases.

Overview

What this was about

Kenzo Milleville presented a method to convert pixel measurements on specimen images into metric units by reading the ruler on the sheet. Instead of predicting scale directly, the ruler is detected and normalised, similar rulers are retrieved from an annotated reference set with a CLIP-like similarity model, and a keypoint-matching model transforms the known tick marks onto the new image, requiring in principle only one labelled example per ruler type. The reference set grew to about 2,000 rulers (21,000 annotated points, ~350 institutions), and 96% of predictions now have under 1% error (mean ~0.3%, below the ~0.8% manual annotation error). A proof of concept with Segment Anything 3 measured leaf sizes across taxa and geography. He recommended including ArUco markers or QR codes of known size during digitisation, and noted a proposal to Audiovisual Core to standardise such metadata and the idea of sharing length annotations via DiSSCo.

Why it matters. Normalised measurements are a prerequisite for comparing traits across collections and at global scale from digitised specimens.

Key ideas

In the room

  • Rulers are highly diverse but frequently reused across institutions.
  • Pipeline: object detection, normalisation, similarity retrieval (top 3), keypoint matching, post-processing.
  • Main improvement over last year's Living Data version comes from a larger dataset and renewed models, especially keypoint matching.
  • Robust to lower resolution and partial occlusion; high-contrast popular rulers perform best; simple block rulers worst.
  • Annotation error is roughly constant in pixels regardless of image resolution.
  • No standard yet for storing calibration metadata; proposal submitted on GitHub for Audiovisual Core.
Jump in

Notable moments

In their words

Transcript

Automatically generated captions can contain mistakes, especially in names and technical terms. Times are relative to the room recording.

Read the transcript ↓
Loading transcript…