LLM-Based Pipeline for Extracting Nomenclatural Acts from Taxonomic Literature
Ishaipiriyan Karunakularatnam · Large Language Models for Biodiversity Data Discovery, Integration, and Curation. Part A: AI- enabled knowledge creation and extraction
The short versionA grounded, schema-constrained two-pass LLM pipeline extracts IPNI-ready nomenclatural data well once treatments are found; finding treatment boundaries is the bottleneck.
What this was about
Ishaipiriyan Karunakularatnam presented a two-stage LLM pipeline to help IPNI curators extract nomenclatural acts from taxonomic PDFs. PDFs are converted to Markdown (PyMuPDF), cleaned of tables and images, and enriched with Crossref metadata. An LLM proposes treatment anchors, which are kept only if they match real document text (exact or ≥90% fuzzy match), and each treatment is then extracted separately against a 30+ field schema. On 50 treatments, anchor detection (F1 0.79) was the weakest link, while name fields were near the ceiling. The model-agnostic LangChain pipeline (tested with a Qwen 27B model via vLLM on an HPC) sits behind a Gradio interface for curator review.
Why it matters. IPNI feeds global backbones like GBIF, so errors propagate downstream. The pipeline shows how to add LLM extraction to high-accuracy curation while keeping curators in control.
In the room
- IPNI holds over 1.4 million plant names from 18,500 publications by 55,000+ authors, curated by Kew's Plant and Fungal Names team.
- Treatment text is abbreviation-heavy, partly Latin and inconsistently formatted, so rule-based parsing does not generalise; curation is manual and slow.
- Pre-processing: PyMuPDF PDF-to-Markdown, removal of images, captions and tables, Crossref metadata with fallback defaults.
- Anchor detection: the LLM proposes treatment starts, which are grounded against document lines by exact then fuzzy matching (90% similarity threshold) to prevent hallucinated anchors.
- Each treatment is extracted separately to keep context small, against a finite schema of 30+ fields (taxonomy, basionym, types, collection, geography, optional descriptions). Failed validations are logged without stopping the paper.
- Model-agnostic through LangChain (Hugging Face transformers locally or vLLM on HPC). Small 7B models hallucinated more and were unreliable on long treatments.
- Preliminary evaluation on 50 treatments: anchor detection F1 0.79; name fields near the ceiling conditional on detection.
Notable moments
Transcript
Automatically generated captions can contain mistakes, especially in names and technical terms. Times are relative to the room recording.


