The GBIF data index: crawling and crawled
Matthew Blissett · Building resilient data infrastructures for biodiversity science
The short versionAI-era crawling through residential proxies is a new, hard-to-block load that is forcing open biodiversity infrastructures and their publishers to spend effort on defence.
What this was about
Matthew Blissett traces how GBIF moved from distributed DiGIR/TAPIR/BioCASe queries to a central occurrence index (2007) and now weekly Darwin Core Archive downloads and image caching, and then describes how usage patterns changed in the last two to three years. Search engines and self-identifying LLM crawlers can be rate-limited or blocked (one that ignores robots.txt was blocked), but 'vibe-coded' apps now hammer the GBIF API with millions of inefficient requests, and residential-proxy networks such as Bright Data send one or two requests per IP while posing as browsers. Cloudflare blocks less than half; GBIF has had to simplify and cache web pages, categorize traffic and push monthly cloud exports, while resisting API keys.
Why it matters. It documents a concrete, growing operational threat to open biodiversity data services and to small publishers whose public image URLs are scraped directly.
In the room
- Early 2000s DiGIR/TAPIR/BioCASe fan-out searches were unreliable; from 2007 GBIF built a central index, crawling publishers with a distributed lock of one request per server.
- Now nearly all publishing uses Darwin Core Archives or data packages; GBIF checks weekly and respects HTTP Not Modified.
- GBIF caches publisher images for at least a year; from early spring this year publishers were overwhelmed by probable LLM-training crawlers fetching public image URLs directly.
- LLM services open thousands of simultaneous connections, unlike past crawlers.
- Search engines and identifying LLM crawlers respecting robots.txt are manageable; one major crawler ignoring it was blocked outright.
- Since March 2025 vibe-coded apps (e.g. "Bird.world", "BioTag") make inefficient API use; over 100 sites made more than 100,000 requests in recent months.
- Residential proxy networks (e.g. Bright Data via apps on smart TVs) make crawling almost unidentifiable; Cloudflare blocks less than half.
Notable moments
Transcript
Automatically generated captions can contain mistakes, especially in names and technical terms. Times are relative to the room recording.


