Machine-Ready Botanical Data: Mapping Integrated Collections and Traits with Darwin Core Data Package
Nathan VanderKraats · Responsible AI, Open Digital Curation, and Round-Tripping for Biodiversity Data
The short versionDarwin Core Data Package can serve as a semantic integration layer for a botanical garden's diverse collections and trait data, and MBG's draft mapping is offered for community refinement.
What this was about
Nathan VanderKraats presented a draft mapping of Missouri Botanical Garden's collections to the Darwin Core Data Package, used as an integration layer across independently developed living collections, herbarium and trait data. He highlighted the value of a semantic graph (queries that work whether or not a system has Tropicos' intermediate 'collection' level) and walked through living collections, herbarium identifications and a detailed specimen traits section distinguishing material assertions, media, regions of interest, traits and ML features with human and AI agents. He argued that capturing which traits drove an identification is a separate problem outside the data package, and invited community input to turn the draft into a botanical profile.
Why it matters. Gives a concrete, institution-scale application of the newly ratified Darwin Core Data Package, including AI-relevant trait and media modelling.
In the room
- The mapping is explicitly a draft; he asked the community to help make it a firmer botanical standard.
- Boxes show both the official DwC-DP class and a convenience label (e.g. 'specimen' for Material).
- Tropicos has an extra 'collection' layer (duplicates from one plant/population) between collection event and specimen; graph queries can answer 'how many specimens from this event?' regardless.
- DwC-DP as an integration layer across independently developed databases, with a warehouse-like relational layer on top; knowledge graphs can improve LLM querying of large databases.
- Some constraints must be enforced externally, e.g. accession numbers shared between organism and seed material, and only one accepted identification per occurrence.
- Identifications link to materials and media (specimen, DNA sequence, photograph) as evidence.
- Traits section: material assertions by humans on physical specimens; media file vs media content; regions of interest and transformations; traits vs ML features; human and AI agents; hyperspectral classes (see Matt Austin's talk).
Notable moments
Transcript
Automatically generated captions can contain mistakes, especially in names and technical terms. Times are relative to the room recording.


