Validating LLM-Extracted Biodiversity Data to Preserve Scientific Meaning
Brooke Long-Fox · Large Language Models for Biodiversity Data Discovery, Integration, and Curation. AI applications, integration, and validation
The short versionLLM extraction speeds up morphological matrix curation enormously, but errors, including fabrications, appear in almost every output, so expert validation is essential.
What this was about
Brooke Long-Fox explained how MorphoBank used an LLM tool, developed with intern Shreya Jariwala over about a year, to restore missing character and character-state definitions to ~600 NSF-funded phylogenetic matrices from PBDB. The tool is now a beta feature for users. It cut per-matrix curation from hours to minutes. Across 400+ papers and ~35,000 character-state entries accuracy was about 90%, but almost every matrix had some error: omissions, misaligned states, fabricated runs of characters, and specialist anatomical terms replaced with common words. Her message was that AI produces a draft and human validation keeps the data scientifically meaningful.
Why it matters. Data can look FAIR while silently changing an author's scientific meaning. Error types from real curation show why repositories must build in human validation.
In the room
- MorphoBank holds 1,500+ phylogenetic morphological character matrices from thousands of published papers.
- Matrices are taxa x characters scored 0/1/2 etc.; the meaning lives in character and state definitions.
- NSF funding supported adding ~600 PBDB Nexus files; many lacked state definitions, which took 1-2 to 10+ hours per matrix to add by hand.
- The AI tool was trained on 20-50 hand-curated datasets; it needed page hints and expected character counts to find the right tables.
- Now a beta feature in MorphoBank, with a warning that AI-extracted characters may contain errors (preprint on bioRxiv).
- Results: 400+ papers, ~35,000 character-state entries, ~90% estimated accuracy; about three hours of work reduced to three minutes.
- Error types: omissions, dense/irregular tables, heading misalignments, fabricated states, ambiguous anatomical terms resolved wrongly.
Notable moments
Transcript
Automatically generated captions can contain mistakes, especially in names and technical terms. Times are relative to the room recording.


