Integrating Many Data Sources in iChatBio Enables Emergent Scientific Use Cases
Michael Elliott · Large Language Models for Biodiversity Data Discovery, Integration, and Curation. AI applications, integration, and validation
The short versioniChatBio's specialised agents let a conversational interface combine many biodiversity data sources, with traceable, non-hallucinated data artefacts.
What this was about
Michael Elliott presented iChatBio, a public multi-agent AI system for biodiversity data. A chat agent delegates sub-tasks to specialised agents for GBIF, iDigBio, iNaturalist, ALA, OBIS, WoRMS, BHL, NCBI, GloBI, NASA climate data, web search and data processing. Worked examples combined GloBI and GBIF to find predators of the white-throated dipper with records in Norway, and generated descriptions and categories for specimen images. He stressed that chat text can hallucinate but data artefacts are real and traceable, and that more agents create emergent use cases.
Why it matters. It lowers the barrier to multi-source questions no single portal can answer, and its open SDK invites the community to add data sources.
In the room
- Public since February 2025 at ichatbio.org; by May 2026 over 2,600 user conversations, 14 agents, 11 data sources, 12 student contributors.
- A chat agent interprets requests and delegates to specialised agents; an open-source Python SDK lets others build agents.
- Agents cover GBIF, iDigBio, iNaturalist, ALA, OBIS, WoRMS, BHL, NASA POWER, NCBI, Google web search, GloBI, plus processing and map generation.
- Vision: data access, processing (filtering, enrichment, statistics, visualisation, FAIR packaging) and inference (Q&A, intelligent cleaning, hypothesis generation), only partly implemented.
- Artefacts record exactly how data were retrieved (search parameters, requests).
- Example: predators of Cinclus cinclus from GloBI, then GBIF counts in Norway; the chatbot looked at only about 10 species to keep responses short.
- Moving toward user-defined repeatable workflows, e.g. describing and categorising specimen images from GBIF.
Notable moments
Transcript
Automatically generated captions can contain mistakes, especially in names and technical terms. Times are relative to the room recording.


