talk · Thursday 24 September · FORUM

Canada’s need for synthesis for Adaptation and Resilience to Climate Change

James Macklin · Sustaining the Geo- and Biodiversity Data Ecosystem: Standards-Driven Approaches to Long-Term Resilience

Recording time 1:24:25–1:36:43Open on Vimeo ↗

The short versionGenomic data can only support climate adaptation synthesis if metadata are captured with standards from the start and linked to discoverable infrastructures like GBIF.

Overview

What this was about

James Macklin drew lessons from GenARCC, a large Canadian cross-departmental project using genomics to assess, predict and adapt to climate change, which will add about 20,000 whole genomes across 30 species largely from existing samples and specimens. He showed how the DINA research data management system captures standards-based metadata linked to material samples and sequences, and new pipelines push data semi-automatically to ENA and then associate ENA records with rich metadata in GBIF, compensating for poor metadata searchability in INSDC repositories. He stressed that relied-upon repositories carry funding risk, and that researchers undervalue the 'boring' work of connecting and standardising metadata needed for reuse and decision-ready results.

Why it matters. Environmental genomic data are growing rapidly, and the talk shows a practical pipeline (DINA to ENA to GBIF) for making sequence data findable while highlighting cultural barriers to reuse.

Key ideas

In the room

  • Public genomic data are estimated at 70+ petabytes; environmental metagenomics is a rapidly growing, highly diverse slice.
  • Canada has genome-level data for about 0.5% of its 80,000 species, biased toward economically and culturally important species.
  • GenARCC will add about 20,000 whole genomes across 30 terrestrial, marine and freshwater species, using existing samples via MTAs.
  • DINA captures a standards-based metadata snapshot with flexible managed attributes for terms not covered by standards.
  • About 740 salmon genomes were pushed from DINA to ENA via a semi-automatic pipeline.
  • INSDC repositories hold only core metadata searchably; a new workflow links ENA records to rich metadata in GBIF.
  • Researchers prioritise novel sequencing and analyses over metadata curation, reproducible code and reusable pipelines.
Jump in

Notable moments

In their words

Transcript

Automatically generated captions can contain mistakes, especially in names and technical terms. Times are relative to the room recording.

Read the transcript ↓
Loading transcript…