talk · Thursday 24 September · FORUM

FAIR ancient environmental DNA data in practice: developing and implementing interoperable metadata standards in AEGIS

Joana Pauperio · FAIR data in practice - eDNA, modelling and biodiversity concepts

Recording time 2:44:21–2:58:47Open on Vimeo ↗

The short versionCommunity-driven ancient DNA metadata standards, now in MIxS, can make paleogenomic data FAIR and reusable across biodiversity infrastructures.

Overview

What this was about

Joana Pauperio described metadata work for the AEGIS ancient environmental genomics programme, in which ancient DNA from sources such as lake sediment cores flows through ENA and MGnify to the AEGIS data portal. The team mapped AEGIS metadata to the ENA data model and to the MInAS community standard, working with the Genomic Standards Consortium and FAIRe, and created an interim ENA ancient DNA checklist of 48 attributes including past environmental context and age inference. With MInAS now incorporated into MIxS version 7, ENA will transition to the MIxS ancient checklist, and experiment metadata work is aligning with a GA4GH core checklist.

Why it matters. Ancient eDNA provides long-term records of ecosystem resilience, and standard metadata (including past environmental context and age) are needed for it to join wider biodiversity data ecosystems.

Key ideas

In the room

  • AEGIS is a seven-year programme using ancient environmental DNA to understand ecosystem resilience and inform sustainable food production.
  • Reference genomes (via the Earth BioGenome Project standards) support interpretation of ancient eDNA.
  • Ancient eDNA metadata map to ENA's study, sample, experiment and run model, including ancient-specific fields such as age estimate and past environmental context.
  • An interim ENA ancient DNA checklist has 48 attributes, 13 mandatory and 17 checklist-specific.
  • MInAS was incorporated into MIxS version 7 this summer; ENA will move to the MIxS ancient checklist.
  • Experiment metadata work with GA4GH yields a core checklist that maps well to INSDC mandatory fields.
  • GBIF already ingests sequence-based data from ENA, and ancient DNA occurrences may flow there too.
Jump in

Notable moments

In their words

Transcript

Automatically generated captions can contain mistakes, especially in names and technical terms. Times are relative to the room recording.

Read the transcript ↓
Loading transcript…