talk · Thursday 24 September · SAL B

AI Ready Standards with Croissant for Type Specimens Catalog Datasets

Sefika Efeoglu · AI-Readiness Metrics and Metadata for Biodiversity Data

Recording time 4:19:54–4:29:24Open on Vimeo ↗

The short versionCombining LLM extraction, Darwin Core and Croissant metadata offers a route from historical specimen catalogues to FAIR, ML-ready datasets.

Overview

What this was about

Sefika Efeoglu presented a workflow for turning historical type specimen catalogues into AI-ready datasets: LLMs (e.g. GPT and Qwen) extract structured information from unstructured catalogue text, which is expressed in Darwin Core and then described with ML Commons Croissant metadata in JSON-LD. She showed how Darwin Core fields map to Croissant fields (scientific name as label, genus as taxonomic feature, individual count as integer, catalogue number as identifier, locality, event date, collector, type status) and how the pipeline evaluates FAIR metrics. The approach aims to bridge biodiversity standards and ML dataset platforms such as Hugging Face, Kaggle and OpenML.

Why it matters. Croissant is becoming the metadata layer ML platforms understand; mapping Darwin Core into it could make collection data directly discoverable and loadable in ML ecosystems.

Key ideas

In the room

  • Biodiversity data are used for species identification, named entity recognition, taxonomic knowledge extraction and knowledge graphs, but historical catalogues are largely unstructured text.
  • FAIRness supports ML reproducibility, clarity, sharing and more reliable model development; the aim is to extend FAIR into ML workflows.
  • Croissant (ML Commons) is a standardised, machine-readable metadata format for ML-ready datasets supported by platforms such as Hugging Face, Kaggle and OpenML: 'create once, publish anywhere, use everywhere'.
  • Workflow: type specimen catalogue, unstructured text, LLM-assisted extraction (GPT, Qwen), Darwin Core metadata, Croissant metadata, AI-ready dataset.
  • Mapping examples: scientificName as label for classification, genus as taxonomic feature, individualCount as integer feature, catalogNumber as unique identifier, locality, eventDate, collector (provenance) and typeStatus.
  • Records are represented in Croissant JSON-LD with explicit field data types; the pipeline measures FAIR metrics and metadata quality.
Jump in

Notable moments

In their words

Transcript

Automatically generated captions can contain mistakes, especially in names and technical terms. Times are relative to the room recording.

Read the transcript ↓
Loading transcript…