talk · Friday 25 September · SAL C

Reporting structure and sampling-bias correction in citizen-science bird data across Asian cities

Yeshan Qiu · From Standards to Implementation: Connecting Observation Data in Asia to GBIF Infrastructure

Recording time 3:16:58–3:33:47Open on Vimeo ↗

The short versionAbundant citizen-science records do not guarantee inference support; reporting structure can distort richness hotspots and urban-biodiversity relationships unless it is diagnosed and controlled.

Overview

What this was about

Yeshan Qiu, from the user side, examined how citizen-science reporting structure shapes what GBIF bird data can support in urban biodiversity analysis. Using about 2.6 million cleaned 2024 GBIF bird records across 13 Asian cities on 2 km grids, she characterised spatial support, environmental representation, inventory/observer support and effects on ecological inference. Reporting was uneven across and within cities, concentrated in accessible places (dense roads, near parks and water) and increasingly compressed in larger cities; standardising to 10 reporting events per grid shrank the apparent richness gap between high- and low-reporting grids from 2.9x to 1.6x and changed urban pattern-biodiversity relationships (road density effect from about -0.2 to -0.4).

Why it matters. Cross-city urban ecology increasingly relies on GBIF and eBird data; this work shows when conclusions are safe and argues for richer sampling-event metadata and fitness-for-use frameworks.

Key ideas

In the room

  • Observed occurrence patterns combine the ecological process and an observation/reporting process driven by where people go and report.
  • Framework of four gaps: spatial reporting support, environmental representation, inventory and observer support, and effect on ecological inference.
  • Data: GBIF bird occurrences from 2024 across 13 Asian cities, about 2.6 million records after cleaning, about 2 million within 2 km analysis grids; predictors cover accessibility, built form, vegetation, water and city-level descriptors.
  • Hong Kong, Mumbai, Seoul and Singapore cover nearly the entire city; Beijing, Dhaka and Shanghai are thin in space and time; within-city structure varies (diffuse, single core, multiple hotspots).
  • Sampled grids have denser roads and buildings and are closer to parks; this selectivity strengthens with city size, compressing reporting into a smaller footprint.
  • Parks are associated with more records, more observers and broader temporal coverage.
  • Standardising to 10 reporting events per grid reduced the apparent richness advantage of high-reporting grids from about 2.9x to about 1.6x.
Jump in

Notable moments

In their words

Transcript

Automatically generated captions can contain mistakes, especially in names and technical terms. Times are relative to the room recording.

Read the transcript ↓
Loading transcript…