AI for Nature at Large Scale and High Resolution
Tanya Berger-Wolf · Monday plenary programme
The short versionAI can help fill biodiversity knowledge shortfalls only if biological knowledge is built into the models and the data infrastructure is FAIR for AI and returns value to primary data collectors, because 'there is no AI without data'.
What this was about
Berger-Wolf argued that computation does not change the scientific method but lets scientists 'look at more things, more carefully', at larger scale and higher resolution. Biodiversity data, however, are heavily biased toward where wealthy people are, and AI methods tuned to benchmarks tend to learn more about what we already know. She presented Imageomics work that builds biological knowledge into machine learning. The examples were BioCLIP (a foundation model trained on 214 million images of about a million species captioned with full taxonomy), biologically meaningful interpretability with attention and sparse autoencoders, automated trait measurement of NEON ground beetles, and drone-based behaviour data. She then showed that AI agents fail mainly at finding and preparing data rather than at analysis. She ended on data readiness: the FAIR for AI checklist and agent, sovereign data supply chains, and infrastructure that returns value to primary data collectors.
Why it matters. The keynote connects state-of-the-art machine learning to TDWG's core business. Its evidence that AI agents fail mostly at data discovery, API construction and data preparation, rather than at analysis, argues that standards, metadata and provenance are the bottleneck for AI in biodiversity science. It also calls for provenance and attribution to flow back from ML-derived products to the original data providers.
In the room
- Following Poincaré's Science and Method, computation lets scientists look at more things more carefully, but data remain strongly skewed geographically, taxonomically and by wealth. Observations track where rich people are, not where biodiversity is.
- Estimated species numbers have risen to roughly 30-50 million while only about two million (perhaps 1.5 million unique) are named, so the Linnean shortfall is growing. A recent prospective paper maps each biodiversity shortfall to AI research that could help fill it.
- 'AI' here means mostly (about 99%) machine learning, not LLMs. Benchmark-driven ML improves performance on the data-rich bulk of skewed distributions rather than the long tail where discovery lies.
- BioCLIP was trained on about 214 million images of about a million species (sources include EOL, GBIF, iNaturalist, FathomNet and BIOSCAN), simply captioned with full taxonomic strings. It learns taxonomy implicitly and can classify at any rank.
- BioCLIP's embedding space holds biology it was never given: clusters for life stage, sex and disease, a Darwin's finch axis aligned with beak size, and a freshwater versus marine split in fish. Similarity search surfaces diseased plants and pollinator interactions.
- Biologically meaningful interpretability: attention-based part highlighting, and sparse autoencoders that manipulate features to show what about a trait (for example the shape of a house sparrow's beak) distinguishes a species. This supports trait discovery, dichotomous keys and separating mimic species.
- An imaging pipeline for NEON ground-beetle bycatch segments specimens and measures landmark distances more accurately than calipers. The traits are being tested as predictors of drought and soil moisture.
Notable moments
Transcript
Automatically generated captions can contain mistakes, especially in names and technical terms. Times are relative to the room recording.


