talk · Thursday 24 September · SAL B

Provenance, Lineage, and Auditability in AI-Driven Biodiversity Image Workflows

Xiaojun Wang · AI-Readiness Metrics and Metadata for Biodiversity Data

Recording time 3:26:56–3:36:29Open on Vimeo ↗

The short versionTwo persistent-identifier links per derived image, parent and batch, are enough to preserve auditable lineage and pipeline context for AI-processed biodiversity images, even outside the repository.

Overview

What this was about

Xiaojun Wang presented how the Fish-AIR repository records provenance for AI-derived fish images produced mostly by external pipelines. Fish-AIR holds just over 400,000 source images from 67 institutions plus about 306,000 derived images (bounding boxes, segmentations, landmarks), and each derived record carries a parent ARK (its direct input image) and a batch ARK (the processing run), so following parent links reconstructs lineage back to the specimen image. Batch records (a class with 14 terms in the Fish-AIR vocabulary) store shared run metadata once, and ARKs, parent and batch references travel in exported CSV and RDF downloads. Automated checks of about 359,000 parent links and 763,000 batch links found all resolvable, enabling researchers to trace odd results and compare model performance across collections or batches.

Why it matters. As AI pipelines generate derived media at scale, lineage that survives export is essential for reproducibility and for diagnosing whether model errors come from source material or processing.

Key ideas

In the room

  • Fish-AIR serves the Biology-Guided Neural Networks and Imageomics projects; derived outputs are produced mostly by external teams' pipelines.
  • Snapshot: just over 400,000 source image records from 67 institutions and 7 source collections; about 306,000 derived images, including roughly 290,000 bounding box images, 42,000 segmentations and 27,000 landmark images.
  • A specimen image can pass through bounding box, segmentation and landmark tasks; each output is a separate record in a separate batch, and file names cannot preserve relationships.
  • Each record has its own ARK; derived images also store a parent ARK and a batch ARK; source images have no parent, and harvested images reference harvest batches.
  • 21 batch records describe AI pipeline runs by the Tulane and Virginia Tech teams; batch is a class with 14 terms (pipeline, institution, creator, project, date, supplement path, code repository, etc.).
  • The largest batch contains almost a quarter of a million outputs; shared metadata is stored once on the batch record.
  • Downloads include multimedia and batch CSVs with ARKs, citation information, term definitions and an RDF model (Fish-AIR vocabulary version 3).
Jump in

Notable moments

In their words

Transcript

Automatically generated captions can contain mistakes, especially in names and technical terms. Times are relative to the room recording.

Read the transcript ↓
Loading transcript…