talk · Thursday 24 September · SAL C

On-Device AI for Data Cleaning, Standardisation, and Exploration in Collections Management

Jack Hollister, Unidentified co-presenter · From Mobilizing Data to AI-Ready Knowledge: Infrastructure for Multimodal Biodiversity Data

Recording time 4:27:50–4:41:52Open on Vimeo ↗

The short versionLocal LLM hardware can clean and enrich millions of legacy collection records at predictable cost, but validation of the outputs is the unsolved problem.

Overview

What this was about

Jack Hollister and an NHM colleague tested whether local LLMs on borrowed Dell GB10 hardware (two clustered units, models up to ~400B parameters) could clean legacy collections data. They held out about 10% of records as test data. Using NVIDIA Nemotron Super with a custom prompt and harness, 24-hour runs recovered missing countries for over 4,000 site records (about 80% correct). Title recovery for bibliography records was hampered by websites blocking LLM scraping. Georeferencing about 6,500 localities gave a median error of about 5.3 km. Scaling to the 6.4 million DiSSCo UK records lacking coordinates would take over 1,000 days on GB10s. A GB300 workstation (~£100,000) ran 100 concurrent models to process a deduplicated 4.2 million records in 9.1 days (median ~12 km). They stressed that validation, side-by-side storage of enrichments and local 'token economics' are the key open issues.

Why it matters. Museums hold large amounts of incomplete legacy data that are sensitive or costly to send to cloud APIs. On-premises AI with a safe institutional sandbox offers a different cost and governance model.

Key ideas

In the room

  • The idea is to bring AI onto local hardware rather than the cloud, giving researchers a safe environment to prototype. NHM IT initially refused agents, which led the museum to build a safer isolated infrastructure.
  • Three 24-hour experiments on full data loads: sites (~700,000 records, 58 columns, recover missing country), bibliography (missing titles, 113 fields, ~60,000 records), GPS (6.2 million records, 4.4 million without usable coordinates).
  • About 10% of records with known answers were held out as test data.
  • They tested Qwen and open GPT models and chose NVIDIA Nemotron Super. The prompt defines the task and allowed evidence; the harness prepares data, calls the model and checks output. 'The LLM is only a small part of the system.'
  • Country recovery: over 4,000 records in 24 h, ~80% exactly correct on test data, ambiguous cases rejected; the full set would take 46 days.
  • Bibliography: 2,300 records examined, 66 duplicates found, but only 125 titles recovered because websites block LLM scraping.
  • Georeferencing: ~6,500 records in 24 h with ~5.3 km median error. DiSSCo UK has ~6.4 million records with verbatim localities but no coordinates, which would need over 1,000 days on GB10s.
Jump in

Notable moments

In their words

Transcript

Automatically generated captions can contain mistakes, especially in names and technical terms. Times are relative to the room recording.

Read the transcript ↓
Loading transcript…