A conference theme

LLM evaluation and hallucination

Measuring the accuracy, uncertainty and cost of LLM outputs, and detecting or controlling hallucination.

12 talks & discussions
Across rooms and sessions

Explore the conversation.

talk · Tuesday 22 September · SAL A

AI-assisted digitisation of printed identification keys: from book to structured data in minutes

An agentic LLM pipeline can turn printed keys and descriptions into a structured draft identification key for tens of dollars, shifting experts from transcribers to reviewers.

Wouter Koch ↗
talk · Tuesday 22 September · SAL A

HerbAudit: Validating AI-Driven Herbarium Transcriptions

HerbAudit shows AI herbarium transcription reaching ~94% accuracy and provides a fair, field-aware way to benchmark it.

Dilara Ağacık ↗
talk · Tuesday 22 September · SAL A

Introduction to the AI for Biodiversity Data session: what machine learning and LLMs actually do

LLMs are word predictors with no notion of truth, so their outputs, including hallucinations, cannot be objectively scored from the text alone.

David Williamson ↗
talk · Tuesday 22 September · SAL B

KakraCards: An AI-Assisted Pipeline for Liberating Six Decades of Seabird Heritage Data

Multi-model consensus with human adjudication builds ground truth and picks the best LLM for transcribing standardised historical cards.

Kristjan Adojaan ↗
discussion · Tuesday 22 September · SAL B

LT17 closing Q&A: AI hallucinations, validation and lightning-talk format

Validation of large-scale AI output remains open; for ML classifiers, accuracy should be judged on the ecological task rather than benchmark metrics.

Jack Hollister, Haris, Tanya Berger-Wolf ↗
talk · Tuesday 22 September · SAL A

Multimodality for Knowledge Extraction from Historical Entomology Literature

Aligning text and illustrations in a multimodal knowledge graph can generate evidence-grounded captions and recover knowledge lost by text-only extraction.

Jana Hoffmann, Sefika Efeoglu ↗
talk · Thursday 24 September · SAL C

A metadata-first FAIR4AI checklist and automated evaluation agent for biodiversity datasets

FAIR4AI is mainly a metadata gap, and a community checklist plus an automated agent can measure and help close it while tracking provenance and governance back to data sources.

Tanya Berger-Wolf ↗
talk · Thursday 24 September · SAL B

From 'Lost in Translation' to Fit-for-Purpose Data: Why Ecologists and Conservation Practice Still Need Groupings and Uncertainties

Biodiversity data systems must be able to record identification uncertainty and taxonomic groupings, or ecologists face a choice between false certainty and lost ecological information.

Bernhard Kløw Askedalen, Natalie ↗
talk · Thursday 24 September · SAL C

Mobilising and integrating extracted information from the literature into a biodiversity data ecosystem - the BIOfid approach

BIOfid mobilises Central European biodiversity literature with NLP, ontologies and standards, because raw LLM extraction still hallucinates identifiers and needs verification.

Gerwin Kasperek ↗
talk · Friday 25 September · SAL A

From Historical Oology Cards to Structured Biodiversity Data: Multimodal LLM Transcription, Uncertainty Metrics, and Human Review

Token-level hesitation from log-probabilities does not tell you which transcriptions are wrong, but it reliably predicts where human review effort will go.

Grete Pasch ↗
talk · Friday 25 September · SAL A

LLM-Based Pipeline for Extracting Nomenclatural Acts from Taxonomic Literature

A grounded, schema-constrained two-pass LLM pipeline extracts IPNI-ready nomenclatural data well once treatments are found; finding treatment boundaries is the bottleneck.

Ishaipiriyan Karunakularatnam ↗
talk · Friday 25 September · SAL A

Validating LLM-Extracted Biodiversity Data to Preserve Scientific Meaning

LLM extraction speeds up morphological matrix curation enormously, but errors, including fabrications, appear in almost every output, so expert validation is essential.

Brooke Long-Fox ↗