discussion · Tuesday 22 September · SAL B

SYM23 panel discussion on restricted access species data

Cameron Slatyer, Dimitri Brosens, Will Morris, Johan Liljeblad, Clara Baringo Fonseca, Charlotte Lassaline, Tania Laity, André Zerger, Arthur Chapman, Ruben Pérez, Dmitry Dmitriev, Niels Raes, Franziska Schuster, Laura Russell, Sylvia Orli, Shelley, Mora Aronsson · To see or not to see - that is the question. Managing restricted access data (RASD)

Recording time 41:04–2:10:38Open on Vimeo ↗

The short versionRestricting data is costly and imperfect; panellists argued for targeted, 'shrink-wrapped', well-justified restrictions handled centrally rather than blanket redaction.

Overview

What this was about

Two panel Q&A sessions chaired by Cameron Slatyer: an impromptu one on the first four talks when the session ran ahead of schedule, and a longer one with all speakers after the final talk. Topics included the real cost of managing sensitivities (several FTE in Sweden, monthly list updates at the ALA, opportunity cost), breaches (none known), personal privacy of collectors, whether generative AI already exposes locations, obfuscation on request of data collectors (e.g. pest hunting and Indigenous datasets), taxonomic upscaling, names that reveal localities, correlated fields such as collector and date, historical data, informationWithheld usage, and the feasibility for global collections like the Smithsonian. It closed with the argument that much RASD reflects expert preciousness rather than real threat and that restriction should be justified formally.

Why it matters. Captures practitioners' candid views on the costs, risks and unintended consequences of RASD, including the conservation harm of over-restriction and the new threat of AI mining.

Key ideas

In the room

  • Cost: Sweden has about three FTE handling access contracts; FinBIF sees a large hidden human cost and opportunity cost; the ALA updates RASD lists monthly but estimates about a million records came in only because RASD treatment was available.
  • No panellist knew of data breaches; Will Morris argued habitat loss, not knowledge of locations, is the main threat and 'security by obscurity' is a poor conservation tool.
  • Collector privacy (Kew question): Sweden requires contactable reporters but allows hiding observations; INBO sometimes attributes data to a team rather than individuals.
  • Generative AI (Gemini, ChatGPT) could already reveal where the Kangaroo Island spider lives because the data were mined years ago, though a long-restricted species was refused.
  • Entire datasets may be generalised as a condition of sharing, e.g. vertebrate pest data (to deter illegal hunters trespassing) and Indigenous datasets.
  • ALA generalises only records attached at genus level, not all species in a genus; privacy-obscured iNaturalist records not shared with government led to an area with a highly endangered species being cleared.
  • Names can reveal localities (e.g. a species named after Studley Park); the extension allows withholding the taxon name while keeping the record discoverable.
Jump in

Notable moments

In their words

Transcript

Automatically generated captions can contain mistakes, especially in names and technical terms. Times are relative to the room recording.

Read the transcript ↓
Loading transcript…