talk · Thursday 24 September · FORUM

Species trait literature mining to support integrative biodiversity distribution modelling

Robert Waterhouse · Planning the Libroscope: creating research ready biodiversity from scientific publications

Recording time 6:58:36–7:11:51Open on Vimeo ↗

The short versionDemonstrating research use cases like trait-informed distribution modelling is key to persuading funders to support literature mobilisation infrastructure.

Overview

What this was about

Robert Waterhouse presented a Swiss National Science Foundation project using text-mined species traits and genomic population structure to improve species distribution and taxonomic richness models. Manually curated trait catalogues are laudable but unscalable, so the project extracts species-trait-value triples from Plazi/TreatmentBank treatments and newly prioritised publications, grey literature and supplementary material, using NER trained on community gold-standard annotations. Trait layers support joint SDMs (e.g. predator-prey) and genomic SDMs recognise population structure and dark taxa; the aim is a demonstrative use case showing funders the value of literature infrastructure.

Why it matters. Adding traits and genomic structure to SDMs could substantially change predictions, and the project connects literature mining directly to ecological research questions.

Key ideas

In the room

  • Trait information links species concepts to both genomic and literature knowledge.
  • Curated trait databases exist for popular taxa but are not scalable.
  • Target: species-trait-value triples extracted from treatments and other literature.
  • Trait layers enable joint species distribution modelling; genomic structure enables genomic SDMs.
  • Taxonomic name uncertainty and dark taxa affect trait use and modelling.
  • Taxonomic groups range from well-studied to uncertain (e.g. ~400 amphipods).
  • Gaps include undigitised treatments, grey literature such as field guides, and data in supplementary materials.
Jump in

Notable moments

In their words

Transcript

Automatically generated captions can contain mistakes, especially in names and technical terms. Times are relative to the room recording.

Read the transcript ↓
Loading transcript…