Bridging the Gap: Strategies for Integrating Person Identifiers into Collection Databases
Frederik Berger · From Cabinets to Clouds - Bridging the gap between collection- and database management
The short versionWikidata's community-built person data can be used to pre-structure candidate matches, but expert judgement and institutional priorities remain essential for identifying collectors.
What this was about
Frederik Berger set out a still-conceptual strategy for linking person name strings in MfN collection data to persistent identifiers via Wikidata. About 30,000 people in Wikidata are linked to collections via property P11146, around 15,000 have Bionomia IDs and 13,500 are referenced from Entomologists of the World. MfN's own data yields about 800,000 name occurrences, 200,000 distinct strings and roughly 13,000 clusters, suggesting about 15,000 people. Only about 1.6% are currently linked to a PID in the central CMS. He compared prioritising frequent names against complex clusters, and proposed a web names directory that reconciles against the collection-related subset of Wikidata, with experts making the final decision.
Why it matters. Agent disambiguation is a major bottleneck for linking specimens. The quantified gap, prioritisation options and restricted-reconciliation idea are directly reusable by other collection-holding institutions.
In the room
- No specimen exists without an agent; Wikidata can broker links between institutional databases and other collection information.
- About 30,000 Wikidata humans are linked via 'collection items at' (P11146); about 15,000 people with Bionomia IDs and 13,500 from Entomologists of the World are also represented.
- The MfN WikiProject 'MfN Berlin Names' (led by Sabine von Mering) lists about 1,100 people.
- MfN name data: about 800,000 name strings from about 30 tables, deduplicated to about 200,000 and clustered into about 13,000 groups, suggesting about 15,000 people.
- Only about 1.6% of people in the central CMS have an unambiguous PID (about 10% if distributed lists are included).
- About 1,500 agents account for up to 90% of collection objects, so prioritising frequent names maximises impact; tackling complex clusters improves authority data quality instead.
- Next step: a web-based names directory mapping clusters, strings, sub-collections and candidate Wikidata QIDs, reconciling only against people already linked to collections.
Notable moments
Transcript
Automatically generated captions can contain mistakes, especially in names and technical terms. Times are relative to the room recording.


