talk · Friday 25 September · SAL A

Integrating Many Data Sources in iChatBio Enables Emergent Scientific Use Cases

Michael Elliott · Large Language Models for Biodiversity Data Discovery, Integration, and Curation. AI applications, integration, and validation

Recording time 3:35:29–3:48:35Open on Vimeo ↗

The short versioniChatBio's specialised agents let a conversational interface combine many biodiversity data sources, with traceable, non-hallucinated data artefacts.

Overview

What this was about

Michael Elliott presented iChatBio, a public multi-agent AI system for biodiversity data. A chat agent delegates sub-tasks to specialised agents for GBIF, iDigBio, iNaturalist, ALA, OBIS, WoRMS, BHL, NCBI, GloBI, NASA climate data, web search and data processing. Worked examples combined GloBI and GBIF to find predators of the white-throated dipper with records in Norway, and generated descriptions and categories for specimen images. He stressed that chat text can hallucinate but data artefacts are real and traceable, and that more agents create emergent use cases.

Why it matters. It lowers the barrier to multi-source questions no single portal can answer, and its open SDK invites the community to add data sources.

Key ideas

In the room

  • Public since February 2025 at ichatbio.org; by May 2026 over 2,600 user conversations, 14 agents, 11 data sources, 12 student contributors.
  • A chat agent interprets requests and delegates to specialised agents; an open-source Python SDK lets others build agents.
  • Agents cover GBIF, iDigBio, iNaturalist, ALA, OBIS, WoRMS, BHL, NASA POWER, NCBI, Google web search, GloBI, plus processing and map generation.
  • Vision: data access, processing (filtering, enrichment, statistics, visualisation, FAIR packaging) and inference (Q&A, intelligent cleaning, hypothesis generation), only partly implemented.
  • Artefacts record exactly how data were retrieved (search parameters, requests).
  • Example: predators of Cinclus cinclus from GloBI, then GBIF counts in Norway; the chatbot looked at only about 10 species to keep responses short.
  • Moving toward user-defined repeatable workflows, e.g. describing and categorising specimen images from GBIF.
Jump in

Notable moments

In their words

Transcript

Automatically generated captions can contain mistakes, especially in names and technical terms. Times are relative to the room recording.

Read the transcript ↓
Loading transcript…