Agentic AI Systems for Specimen Interaction and Camera-Based Observation
Jack Hollister · Bots, Bits, and Biodiversity
The short versionOff-the-shelf local vision language models can pick out and count specimen morphology and spot damage with minimal guidance.
What this was about
Jack Hollister tested whether an off-the-shelf, locally hosted 32-billion-parameter vision language model, with only a prompt and harness, could identify morphological features of insect specimens. Using a 3D-printed rotating scanner driven by a Raspberry Pi to take 40 images per specimen and Dell GB10 compute, the model drew bounding boxes and counted wings, legs and antennae, reported colour, and detected damage; features were stored in an interactive graph network.
Why it matters. Shows a low-guidance, locally hosted route to extracting morphological data and condition information from collections.
In the room
- Local hosting on Dell GB10 machines (~£7,000) which can run up to 200B-parameter language models (400B clustered); the largest usable VLM was 32B.
- No fine-tuning; only a prompt and harness describing tasks and inputs.
- 3D-printed scanner rotates specimens 360° and tilts the camera ~150°, controlled by a Raspberry Pi; 40 images per specimen.
- Asked simply 'what can you see', the model counted and boxed wings, antennae, legs and pins and identified damage.
- Features stored in a graph network for interactive querying.
Notable moments
Transcript
Automatically generated captions can contain mistakes, especially in names and technical terms. Times are relative to the room recording.


