discussion · Tuesday 22 September · SAL A

SYM08B discussion: extracting geography and articles from BHL

Nicole Kearney, Roderic Page · BHL at 20: A New Chapter for Biodiversity Literature and Data - Transforming BHL: Data Extraction, Semantic Navigation and Future Interfaces

Recording time 8:19:39–8:27:26Open on Vimeo ↗

The short versionArticle-level discoverability and geographic extraction are key BHL gaps that AI may now make tractable, but resourcing remains the constraint.

Overview

What this was about

The closing discussion began with an audience provocation: if GBIF and BHL are both in RDF, could an LLM use GBIF as a gazetteer to extract geographic information from all BHL pages, including the old BHL field journals project whose expedition notes could form rich information graphs? A second participant asked why BHL still lacks article-level indexing. Kearney explained content is uploaded at volume level; BioStor and later the Persistent Identifier Working Group added article data and DOIs (about 450,000 articles), but the paid work was paused during the Smithsonian transition, and she urged contributors to supply article metadata. Page added that articles could be found with a complete bibliography of life or by having AI read BHL for article pages, as NHM's work does. Kearney closed by noting BHL's need for fundraising and new ideas.

Why it matters. Highlights practical priorities for BHL (articles, places, field notes) and the link between discoverability and sustainability.

Key ideas

In the room

  • Idea: use GBIF as a gazetteer with LLMs to extract localities across BHL.
  • BHL field notes project: expedition journals could yield rich information graphs.
  • About 450,000 journal articles in BHL have article data, enabling discoverability and DOIs.
  • Paid article-data work was put on hold in April of the previous year during the transition.
  • Article discovery needs a complete bibliography of life or AI reading BHL pages.
  • BHL is new to fundraising and philanthropy.
Jump in

Notable moments

In their words

Transcript

Automatically generated captions can contain mistakes, especially in names and technical terms. Times are relative to the room recording.

Read the transcript ↓
Loading transcript…