talk · Tuesday 22 September · SAL B

Agentic AI Systems for Specimen Interaction and Camera-Based Observation

Jack Hollister · Bots, Bits, and Biodiversity

Recording time 7:25:43–7:30:12Open on Vimeo ↗

The short versionOff-the-shelf local vision language models can pick out and count specimen morphology and spot damage with minimal guidance.

Overview

What this was about

Jack Hollister tested whether an off-the-shelf, locally hosted 32-billion-parameter vision language model, with only a prompt and harness, could identify morphological features of insect specimens. Using a 3D-printed rotating scanner driven by a Raspberry Pi to take 40 images per specimen and Dell GB10 compute, the model drew bounding boxes and counted wings, legs and antennae, reported colour, and detected damage; features were stored in an interactive graph network.

Why it matters. Shows a low-guidance, locally hosted route to extracting morphological data and condition information from collections.

Key ideas

In the room

  • Local hosting on Dell GB10 machines (~£7,000) which can run up to 200B-parameter language models (400B clustered); the largest usable VLM was 32B.
  • No fine-tuning; only a prompt and harness describing tasks and inputs.
  • 3D-printed scanner rotates specimens 360° and tilts the camera ~150°, controlled by a Raspberry Pi; 40 images per specimen.
  • Asked simply 'what can you see', the model counted and boxed wings, antennae, legs and pins and identified damage.
  • Features stored in a graph network for interactive querying.
Jump in

Notable moments

In their words

Transcript

Automatically generated captions can contain mistakes, especially in names and technical terms. Times are relative to the room recording.

Read the transcript ↓
Loading transcript…