Assessing record quality and completeness in geoscience collections
Adam Mansur · State of the Collections Description Interest Group. What's Happening and What's Not!
The short versionElement-level views of MIDS information elements are more useful than raw scores for targeting data improvement in large geoscience collections.
What this was about
Adam Mansur compared the Smithsonian's internal Collection Digitization Reporting System (CDRS) with MIDS for assessing record completeness in the Mineral Sciences (about 500,000 records, largely digitised) and Paleobiology (about 800,000 records against an estimated 10-13 million needed) collections. CDRS rules vary by department, are conditional and opaque, and barely move with large projects, whereas MIDS scores are consistent but cluster tightly around MIDS 1 because of the large jump required to reach MIDS 2. He argued that element-level breakdowns (e.g. via the MIDS calculator) are more useful for targeting work: Mineral Sciences lacks stratigraphic context in most records regardless of acquisition decade, while Paleobiology has gaps in names and object type concentrated in particular collections.
Why it matters. Offers practical guidance for using MIDS to plan digitisation and enhancement work, and highlights needs such as signalling checked-but-null values.
In the room
- NMNH holds an estimated 140 million objects across seven departments, two geology-focused.
- Geology metadata resembles biological metadata plus geologic context (unit, age).
- CDRS measures completeness against a 'standard record' by presence/absence checks per primary collection.
- Departmental CDRS rules are inconsistent, conditional and hard to interpret; big projects barely change collection-wide scores.
- Departments mostly treat media as an enhancement, unlike MIDS; CDRS also includes internal elements like storage location and accession number.
- MIDS scores cluster around level 1, so they are only sensitive to large changes.
- Stratigraphic context is present in only 5-20% of Mineral Sciences records across acquisition decades, an intrinsic feature of incoming data.
Notable moments
Transcript
Automatically generated captions can contain mistakes, especially in names and technical terms. Times are relative to the room recording.


