Soundscape China: Listening to the Sounds of Nature with AI
Congtian Lin · AI for Biodiversity Data
The short versionSoundscape China is building a nationwide acoustic sensor network, database and AI models, gathering 3.6 million recordings in its first year.
What this was about
Congtian Lin presented Soundscape China, a national initiative to build a passive acoustic monitoring network, a nature-sound database and AI models, while involving citizen scientists and sound artists. The project tests sound sensors from different companies against standards (4G/5G data transfer, solar power, over 50 days' storage); about 280 devices have been deployed by over 110 participating organisations at more than 260 sites, yielding over 3.6 million records in one year. The platform includes a citizen-science recording app, an annotation tool with volunteers, a localised extension of Audubon Core, and a self-supervised pre-trained spectrogram model for sound event detection and classification, illustrated with bird-call activity patterns at three sites in Shandong Province.
Why it matters. Standardised, large-scale acoustic data and open datasets are prerequisites for training reliable sound-recognition models and for using soundscapes as biodiversity indicators.
In the room
- Soundscape ecology treats sound diversity as an indicator of biodiversity.
- Framework: platform of sensors, models and database; services for data sharing, annotation, analysis and model training; audiences in monitoring, sound art and citizen science.
- Device standards: 4G/5G automatic transfer, solar-powered, >50 days storage, plus plug-in microphones for citizen scientists.
- About 280 devices, 110+ organisations, 260+ sites.
- Audubon Core was used as the data standard but extended and localised, e.g. for equipment information.
- Pre-training on time and frequency patches of spectrograms supports downstream sound event detection and classification.
- The programme is endorsed by ISSP within the UNESCO science decade.
Notable moments
Transcript
Automatically generated captions can contain mistakes, especially in names and technical terms. Times are relative to the room recording.


