talk · Tuesday 22 September · SAL C

Pinned insect digitization conveyor: image and data transcription workflows

Jessica Bird, Sylvia Orli · Digitizing Collections & Legacy Data

Recording time 5:48:07–6:02:02Open on Vimeo ↗

The short versionAn industrial conveyor workflow with QR-encoded metadata and automated skeletal records makes mass digitisation of pinned insects feasible.

Overview

What this was about

Jessica Bird described the Smithsonian NMNH Entomology conveyor project digitising pinned pollinators (bees, beetles, butterflies, moths, flies), funded for an estimated 360,000 specimens, in a collection of about 35 million specimens that would take 20 people 67 years to digitise traditionally. Pre-curation vets taxonomy in EMu and prints unit-tray header labels with QR-embedded taxonomy and identifiers; on the conveyor, labels are staged, specimens barcoded and three focus-stacked views plus label images captured, with metadata in JSON. Images pass QC in the open-source OSPRY tool, enter the DAMS, generate skeletal EMu records published to GBIF via IPT, while label images feed a separate contracted human transcription workflow.

Why it matters. Demonstrates throughput, cost and QC practices for mass digitisation of one of the hardest specimen types, relevant to anyone planning large entomology projects.

Key ideas

In the room

  • USNM Entomology is the second-largest entomology collection (about 35 million specimens, over 21 million pinned).
  • Project started 2025 on pinned pollinators, funded for about 360,000 specimens.
  • Header labels with QR-embedded taxonomy and unit tray identifiers streamline later record creation; drawer IDs name image folders.
  • Silhouette determines imaging sphere/resolution; three focus-stacked views (about 30 slices each) and a composite label image are taken.
  • QC in OSPRY samples 10-40% of images per folder with standardised error phrasing.
  • Skeletal records are created automatically in EMu and published to the collections website and GBIF.
  • Label orientation could not be automatically corrected, so label transcription was split into its own contracted human workflow.
Jump in

Notable moments

In their words

Transcript

Automatically generated captions can contain mistakes, especially in names and technical terms. Times are relative to the room recording.

Read the transcript ↓
Loading transcript…