The BMD Cubing Engine: Harmonizing Biodiversity and Earth Observation Data
Mathias Dillen · Operationalizing Biodiversity Digital Twins within Data Space Ecosystems
The short versionBMD's cubing engine turns a declarative YAML recipe into reproducible, provenance-documented data cubes that harmonise GBIF occurrences with climate and Earth observation data on one grid.
What this was about
Mathias Dillen presented the cubing engine built in the Biodiversity Meets Data (BMD) project as the central harmonisation layer between BMD's data catalogue and analytical tools for Natura 2000 site managers. He described a modular Python back end (base grid class, raster/vector engines, a data cube orchestrator with one child class per data source) and a slim front end that takes a YAML 'recipe' and returns harmonised raster and vector files on a common master grid, plus a STAC description and a provenance record. He covered supported sources (GBIF, CHELSA climate data, Earth observation data), GBIF-specific options including Monte Carlo resampling of occurrence uncertainty, and plans to add sources, publish a package and deploy the engine as an API service within BMD.
Why it matters. Harmonising heterogeneous biotic and abiotic data reproducibly is a precondition for decision-support tools; a recipe-driven, extensible engine with STAC and provenance output makes these cubes discoverable and reusable rather than one-off analyses.
In the room
- The cubing engine is BMD's central harmonisation layer, feeding harmonised data via an API to analytical tools aimed mainly at Natura 2000 site managers.
- It builds on the B3 project's GBIF SQL API 'cube' functionality, which resamples occurrence data and makes heterogeneous data safer to use.
- Back end: a base grid class (currently EEA grid, a global equal-area grid and WGS84), raster and vector engines for pure geospatial transformations, and a data cube orchestrator with abstract raster/vector cube classes; each data source is a child class, so adding a source means writing one class.
- Front end: a single BMDCube class reads a YAML recipe (global, spatial/temporal and source-specific configuration) and runs the harmonisation without users needing to code.
- Output: a data tree of harmonised raster and vector files on the same master grid, a STAC file describing the cube, and a provenance record (runtime, engine and dependency versions, hardware); files can be streamed rather than downloading the whole cube.
- Stack: Python with GDAL, GeoPandas, xarray (central data model), Dask/DuckDB; outputs GeoParquet for vectors and NetCDF or Zarr for rasters.
- Current sources are GBIF, CHELSA climate data and Earth observation data (being moved to direct access to Copernicus S3 storage).
Notable moments
Transcript
Automatically generated captions can contain mistakes, especially in names and technical terms. Times are relative to the room recording.


