talk · Friday 25 September · SAL A

Extracting Structured Biodiversity Knowledge from Literature with AI-Assisted Workflows

Yağmur Güleç · Large Language Models for Biodiversity Data Discovery, Integration, and Curation. Part A: AI- enabled knowledge creation and extraction

Recording time 2:14:41–2:25:28Open on Vimeo ↗

The short versionA hybrid text-layer/OCR tool can free tables from grey-literature PDFs into CSV with about 83% precision, but it misses a third of tables, especially narrow ones.

Overview

What this was about

Yağmur Güleç presented SLGO's in-house tool for extracting tables from PDFs in grey literature. It uses the text layer for born-digital pages and falls back to Tesseract OCR only for scanned pages. A demo showed PDF upload, table detection with bounding boxes, previews and export of each table to CSV, with merged cells represented as empty cells. On 19 born-digital documents, human validation gave 83% precision and 67% recall. False positives were mostly lists of tables and tables of contents; missed tables were mostly one- or two-column tables.

Why it matters. Much St. Lawrence biodiversity data is trapped in report tables. Automating their extraction could feed ETL pipelines toward standardised datasets for OBIS and GBIF.

Key ideas

In the room

  • SLGO, founded in 2005 with 60+ partners, provides open data on the St. Lawrence ecosystem and is one of three CIOOS regional associations.
  • A large share of biodiversity information sits in grey literature (reports, theses, field notes) and its tables fail the FAIR principles.
  • The tool uses the PDF text layer for born-digital pages and Tesseract OCR only when needed, avoiding costly OCR.
  • Interface: upload a PDF, browse pages, list detected tables with bounding boxes, preview and export each table as CSV.
  • Merged cells are represented as empty cells in the reconstructed CSV.
  • Evaluation on 19 born-digital documents: precision 83%, recall 67% (roughly 75% overall).
  • False positives: lists of tables and tables of contents; misses: mainly one- and two-column tables.
Jump in

Notable moments

In their words

Transcript

Automatically generated captions can contain mistakes, especially in names and technical terms. Times are relative to the room recording.

Read the transcript ↓
Loading transcript…