Literature triage to support Island Biodiversity Monitoring
Patrick Ruch · Planning the Libroscope: creating research ready biodiversity from scientific publications
The short versionClassifier-based literature triage matches human curators at a fraction of LLM cost and is worth investing in for recurring curation tasks.
What this was about
Patrick Ruch presented literature triage: instead of Boolean search, users supply a few hundred relevant papers, from which binary classifiers screen the Biodiversity PMC universe and deliver ranked digests. Evaluations for island biodiversity (BioMoQA) and IPBES assessments showed an ensemble of BERT-type models reaching precision around 90-92%, on par with human inter-rater agreement, far cheaper than generative LLMs; domain-specific biodiversity models did not help. He costed a system at roughly 25K and outlined next steps: ingesting Unpaywall/OpenAlex, LLM-generated regex patterns for post-triage entity extraction (99% species precision), question answering, BioC offset-grounded evidence and downloadable agents.
Why it matters. Scalable, affordable triage helps curators and assessments like IPBES avoid missing relevant literature in rapidly growing volumes.
In the room
- Triage learns from a set of relevant papers and classifies the whole Biodiversity PMC collection, which is larger than PubMed and includes treatments.
- Keep track of negative examples when building triage systems.
- Ensembles of BERT-type models were best and about 3,000 times cheaper than generative chatbots.
- Island biodiversity precision ~90-92%, consistent with inter-rater agreement; IPBES tasks showed even higher agreement.
- Building a system costs roughly one to two person-months plus a few thousand a year hosting, ~25K total; worthwhile for recurring tasks.
- Real LLM costs are subsidised and may rise tenfold; frugal AI is chosen for cost.
- LLMs generate regex patterns from seed keywords for post-triage extraction; a Spanish government pilot extracts species, habitats and fieldwork from environmental assessments.
Notable moments
Transcript
Automatically generated captions can contain mistakes, especially in names and technical terms. Times are relative to the room recording.


