talk · Friday 25 September · SAL A

Contained Agentic Workflow for Literature Analysis and Data Extraction in Museum Collections

Steen Dupont · Large Language Models for Biodiversity Data Discovery, Integration, and Curation. AI applications, integration, and validation

Recording time 3:13:32–3:24:40Open on Vimeo ↗

The short versionContained, local agentic RAG over a literature corpus is feasible on a laptop or shared edge machine, avoiding the security worries of cloud agents in museums.

Overview

What this was about

Steen Dupont demonstrated Kvasir, a prototype built mostly by coding agents (Codex) that runs fully on local hardware. It ingests a personal literature corpus with Docling, stores it in Neo4j and Qdrant, and lets an OpenClaw agent with a local Qwen model answer questions and explore papers. The demo ingested papers into 696 chunks and showed papers broken into sections with their references plotted over time, to characterise papers and trace evidence across the corpus. He argued for local AI for security, noted that optimised local models are improving, and suggested shared edge machines (e.g. a Dell GB10 serving 10-20 users) as institutional infrastructure, with the agent as a pluggable orchestration layer.

Why it matters. Museums worry about releasing agents on institutional data. This prototype shows a low-cost, offline route to experimenting with LLM literature analysis.

Key ideas

In the room

  • The project explores AI on local hardware rather than cloud LLMs; museum IT set up a closed network, but the demo runs offline on a laptop.
  • Motivation: better visibility of a PhD literature corpus: most-cited works, citation timing, references per section, without reading every paper.
  • Architecture: OpenClaw agent orchestrator, Qwen 3.6 main model, Docling parsing, Codex-written Python scripts, Neo4j graph and Qdrant vector store, Kokoro for audio.
  • Demo: two papers ingested, 696 chunks created, a paper view with sections, references per section and their publication years.
  • Idea: the distribution of references (introduction-heavy vs conclusion-heavy) characterises paper type.
  • Why local AI: security, improving local models, and pre-processing so the LLM only does the important step.
  • A GB10-class edge machine could serve 10-20 users on ordinary laptops instead of expensive individual machines.
Jump in

Notable moments

In their words

Transcript

Automatically generated captions can contain mistakes, especially in names and technical terms. Times are relative to the room recording.

Read the transcript ↓
Loading transcript…