talk · Thursday 24 September · SAL A

Old habits and new visions: Exploring the RECODE data model, future integration and data enhancement

Steen Dupont · Solutions for Research Collection Management Systems Challenges

Recording time 8:25:00–8:37:57Open on Vimeo ↗

The short versionNormalising decades of flat CMS data into relational entities and workflow-driven processes gives NHM a foundation for integrations and AI, but it needs institutional governance.

Overview

What this was about

Steen Dupont gave an update on NHM London's RECODE programme (since 2019) to replace a 20-year-old CMS, with go-live expected early next year on the Doxis platform. The data model was heavily normalised: about 70 object record types and 54,000 fields were reduced to about 6,500 in use and then 600–700 fields in four collection object entities (biological, geological, extraterrestrial, human-made). The new system uses true relationships, separates objects from processes, and enforces stricter taxonomy and location hierarchies. Videos of the UAT environment showed a data model viewer, agent relationships, a workflow engine (parallel acquisitions, sequential loans, silent movement control) and a bulk-ingest template builder that outputs CSV/JSON with dynamic errors.

Why it matters. It is a rare look inside a large museum's CMS rebuild, with concrete numbers on field reduction and design choices other institutions can learn from.

Key ideas

In the room

  • RECODE began in 2019; the system is in UAT and expected to go live at the start of next year.
  • Current system: about 70 object record types and about 54,000 fields, reduced to about 6,500 used, then about 600–700 fields.
  • Four target collection object entities: biological, geological, extraterrestrial, human-made.
  • True relationships replace value fields; objects are decoupled from processes; stricter taxonomy and location hierarchies are enforced.
  • Workflow engine supports parallel processes (acquisition), strictly ordered processes (research loans for compliance) and silent processes (movement control with history).
  • A bulk template builder generates CSV and JSON with dynamic error handling for ingest, update and fixes; the JSON can feed other APIs.
  • Business studio admin layer exposes integrations, AI, the data model and search fields; governance is needed to avoid new silos.
Jump in

Notable moments

In their words

Transcript

Automatically generated captions can contain mistakes, especially in names and technical terms. Times are relative to the room recording.

Read the transcript ↓
Loading transcript…