Standardizing Multimodal Insect Monitoring Data for AI-Ready Pipelines: Lessons from an InsectAI Community Datathon
Jamie Alison · AI Driven Monitoring and Species Identification
The short versionCamtrap DP can largely accommodate automated insect monitoring data today, but the community needs agreed guidance and a few targeted extensions (track IDs, parent media and crop offsets, target taxonomic scope, identification provenance) to handle insect-specific challenges.
What this was about
Jamie Alison described an InsectAI community 'datathon' in which participants brought their own insect-camera datasets and tried to map them to the Camtrap DP standard, using a GitHub repository of mini-datasets with automated traffic-light progress reporting and the BioWatch tool as a reward for conformant data. He then set out the data challenges posed by insect sensing systems as a set of 'Cs' (consensus, cropping, counting, coverage and customisation), giving for each a recommendation for what can be done within Camtrap DP today and what might be added beyond it. His overall message was that Camtrap DP works well and should ideally not be changed much, but that the community needs ratified guidance on how to use it for insect data, with longer-term additions such as parent media IDs, track IDs, target taxonomic scope and an assertions table.
Why it matters. Automated insect cameras produce huge, high-temporal-resolution datasets that could offer standardised, relatively unbiased biodiversity measurements, but only if the data can be shared in a common format; the datathon gives concrete, community-derived proposals for how an existing camera-trap standard can be stretched to cover them.
In the room
- A datathon is like a hackathon but focused on converting real data rather than writing code; participants brought stratified 'mini datasets' and conversion code to a shared GitHub repository.
- An automated GitHub traffic-light report on how many standardisation criteria each dataset met motivated participants; BioWatch, which reads Camtrap DP datasets, showed immediate payoff from standardising.
- Consensus: many human and machine agents now identify the same media; within Camtrap DP use the classifiedBy field, and beyond it proposed detections and pipelines (models) tables capture agent provenance; for now avoid multiple determinations per individual within a dataset.
- Cropping: retain media with no observations in the media table (they define temporal scope) and record the parent media ID of crops, e.g. in media comments now, with a future parent media ID and bounding-box offsets in the media table.
- Counting: insects coming and going make Camtrap DP's summable event-based observations hard to apply; recommendation is to use eventID for tracks for now and avoid individualID, with a longer-term 'track ID' concept sitting between event and media-based observation.
- Coverage: Camtrap DP's taxonomic descriptor describes what a dataset contains, not what was searched for; recording target taxonomic scope and invalid samples enables true negatives, possibly by absorbing Humboldt Extension terminology.
- Customisation: diverse monitoring systems should be captured as distinct sampling protocols, and extra data can be added as frictionless resources keyed to existing tables, but best-practice guidance is lacking; an assertions table from the Safe and Sound acoustics project is a candidate.
Notable moments
Transcript
Automatically generated captions can contain mistakes, especially in names and technical terms. Times are relative to the room recording.


