Responsible and GreenAI: Does our community need Gigantic Data Centers to deliver?
Patricia Mergen · Responsible AI, Open Digital Curation, and Round-Tripping for Biodiversity Data
The short versionBiodiversity researchers mostly need AI-ready data rather than giant data centres, and AI use should be weighed against its environmental cost.
What this was about
Patricia Mergen reflected on AI compute needs and trustworthy AI based on her work in the EU AI Factories project. Contrasting SETI@home with today's investment race in data centres and GPUs (US, EU, China), she noted EU sovereignty levels are utopian and trustworthy-AI requirements can block research, while compute dependency disadvantages poorer countries. Although digitising billions of specimens could need millions of GPU hours, researchers she interviewed mostly train on laptops; data cleaning and AI-readiness is the real bottleneck, so she argued for biodiversity thematic data labs and a TDWG role. She closed on environmental costs, including PFAS in data-centre cooling, urging people to think twice before using AI.
Why it matters. Challenges assumptions behind large AI infrastructure investment and points to data standardisation as the community's main AI need.
In the room
- SETI@home-style volunteer computing would be unlikely today due to data scale, cyber-security fears and desire for autonomy.
- Massive investment in data centres is a competition between US, EU and China, with few analyses of actual community needs.
- EU sovereignty certificates (levels 0-4) require fully EU components at the highest level, which is currently unrealistic.
- Trustworthy AI requirements sometimes block research; UNESCO discussions highlight compute dependency of countries without their own infrastructure.
- GPU targets: Amazon claims 2.5 million GPUs in 2025; EU hopes for 1.2 million by 2030; China 1.6 million by 2029.
- ChatGPT estimated analysing 3 billion specimen images might need up to 500 PB and 1-50 million GPU hours, but researchers (e.g. TETTRIs cascade grants) mostly run models on laptops.
- The main time cost is finding, cleaning and making data AI-ready; AI Factories plan thematic data labs, and she is pushing for a biodiversity lab where TDWG standards could play a role.
Notable moments
Transcript
Automatically generated captions can contain mistakes, especially in names and technical terms. Times are relative to the room recording.


