Enabling machine-readable access to biodiversity data through federated semantic querying
El-Amine Mimouni · Designing Institutional Knowledge Data Science Centres for Biodiversity
The short versionDistribute the data, not the infrastructure: cloud Parquet plus OBDA query rewriting gives a virtual Darwin Core knowledge graph without duplicating data into RDF.
What this was about
Mimouni presented the Semantic Data Cloud, a containerised application that exposes distributed Darwin Core Data Package datasets through a single SPARQL endpoint without copying them into a triple store. Datasets are published as cloud-hosted Parquet files with an EML JSON-LD catalogue entry. The system selects relevant datasets by spatial, temporal, licence or maintenance filters, then rewrites SPARQL to SQL through an OWL 2 QL ontology with OBDA mappings based on the Darwin Core conceptual model, and executes it with DuckDB over HTTP. A live instance at data.qcbs.ca holds over 50 GBIF datasets and exposes an MCP server for LLM agents.
Why it matters. It offers a practical route to semantic, cross-dataset querying of the new Darwin Core Data Package without imposing a full RDF stack on data providers, and it is usable directly by LLM agents.
In the room
- The ratified Darwin Core Data Package solves DwC-A's conflation of several entities in one table, but its semantics are implicit in table joins and it is hard to expose online.
- The Darwin Core conceptual model defines relationships in natural language rather than in a formal ontology. An OWL ontology makes the semantics explicit and lets data be seen as a graph (e.g., a material entity evidencing an occurrence during an event).
- Conventional RDF publication costs data duplication, fragile federated queries, a heavy RDF/ETL burden on providers, and schema drift.
- Core principle: 'the data should be distributed, not the infrastructure'. Tables are exploded Parquet files on object storage instead of zipped CSVs.
- An EML JSON-LD file per dataset carries asset hrefs pointing to the Parquet files (as in STAC catalogues) and supports dataset discovery with filters.
- SPARQL is translated to SQL using an OWL 2 QL ontology and OBDA mappings, which guarantees rewritability while remaining a good fit for vocabularies like Darwin Core. DuckDB executes the SQL on remote Parquet over HTTP, so the virtual knowledge graph exists only for the query's duration.
- The layers are separate: storage (Parquet), discovery (metadata catalogue), execution (DuckDB remote SQL), and planning/reasoning (SPARQL plus OWL).
Notable moments
Transcript
Automatically generated captions can contain mistakes, especially in names and technical terms. Times are relative to the room recording.


