AURORA - A Shiny app designed to create Darwin Core Archives from messy biodiversity datasets
Sara Araújo · Semantic Interoperability and Standards Infrastructure
The short versionAURORA lowers the barrier to Darwin Core publication by guiding researchers step by step from messy tables to standard-compliant tables and metadata.
What this was about
Sara Araújo presents the AURORA Shiny app, a browser-based R Shiny workflow of eight steps that turns non-standard occurrence tables into Darwin Core tables ready for publication. It supports tidying, header-to-term mapping with suggestions, creating ID and dependent fields, repository-specific required-term validation (OBIS, GBIF, EMODnet Biology), date and coordinate conversion, taxonomic cleaning and matching against WoRMS or the GBIF backbone, event/occurrence restructuring, eMoF creation, quality control and IPT-style metadata. It does not build the archive itself (the IPT does) and currently supports only a limited set of cores and extensions.
Why it matters. Standardization is a main barrier to data publication; a free, guided tool can bring more datasets into GBIF, OBIS and EMODnet.
In the room
- Eight-step workflow: upload/tidy, map headers to Darwin Core terms, create key fields, validate required terms, transform dates and coordinates, clean and match taxonomy, structure event core/occurrence/eMoF, quality control and metadata export.
- Required terms adapt to the target repository (OBIS, GBIF, EMODnet Biology).
- Taxonomic validation via WoRMS (marine) or GBIF backbone (terrestrial) retrieves scientificNameID and taxonRank.
- eMoF measurements are assigned to events (e.g. temperature) or occurrences (e.g. life stage) and can use controlled vocabularies.
- Quality control produces logs, summaries and issue flags.
- Limitations: does not create the DwC-A (use the GBIF IPT); only event core + occurrence + eMoF or occurrence core + eMoF; users must know their data and basic Darwin Core.
- Funded by the EU via the DTO-BioFlow project; user manual, video tutorial and source code on GitHub.
Notable moments
Transcript
Automatically generated captions can contain mistakes, especially in names and technical terms. Times are relative to the room recording.


