talk · Tuesday 22 September · SAL B

The Edge: Insights into problems that arise from implementing data solutions for restricted access (sensitive) species data

Cameron Slatyer · To see or not to see - that is the question. Managing restricted access data (RASD)

Recording time 53:54–1:04:34Open on Vimeo ↗

The short versionPlan RASD treatments carefully, because taxonomy, data errors and associated records can silently defeat simple rules like 'generalise to 10 km'.

Overview

What this was about

Convener Cameron Slatyer presented the Atlas of Living Australia's experience of edge cases that arise when implementing RASD treatments across more than 184 million records using government-supplied lists. Taxonomy was the biggest problem: undescribed species, disagreements between states, and mismatch between subspecies-level sensitivity and species-level data force precautionary generalisation at higher ranks or whole species. Data errors (wrong coordinates outside a treatment polygon, names or locality in the wrong field, as common with eDNA data) can expose restricted data, and AI mining makes such leaks easier, including an experimental ALA AI tool that exposed a locality recorded in the wrong field.

Why it matters. Practical lessons from one of the largest RASD implementations, directly relevant to anyone automating sensitive-data treatments.

Key ideas

In the room

  • Reasons for RASD range from conservation to discouraging entry to bombing ranges and protecting biosecurity eradication sites for trade reasons.
  • The ALA applies treatments (withholding records and images, generalising, modifying fields) to over 184 million records, driven by lists from state, territory and Commonwealth governments, plus data-provider requests.
  • Main message: plan the approach - what is being restricted, flow-on effects, and risks.
  • Taxonomic problems: undescribed RASD species, disagreement among states over current names, and sensitivity at subspecies level while providers supply species-level data (or vice versa).
  • Rule of thumb: when taxonomic resolution can't be resolved, go up a level; species absent from the backbone attach at genus, so genus-level records get generalised; populations can't be delineated, so the whole species is generalised.
  • Data errors can expose RASD: erroneous coordinates outside a treatment polygon, scientific names in the wrong field (common in eDNA data with BIN numbers), and locality in the wrong field.
  • AI makes such errors easier to mine; an experimental ALA AI tool mining iNaturalist data exposed a sensitive species because locality was in the wrong field.
Jump in

Notable moments

In their words

Transcript

Automatically generated captions can contain mistakes, especially in names and technical terms. Times are relative to the room recording.

Read the transcript ↓
Loading transcript…