talk · Friday 25 September · SAL B

Enabling machine-readable access to biodiversity data through federated semantic querying

El-Amine Mimouni · Designing Institutional Knowledge Data Science Centres for Biodiversity

Recording time 3:26:10–3:35:43Open on Vimeo ↗

The short versionDistribute the data, not the infrastructure: cloud Parquet plus OBDA query rewriting gives a virtual Darwin Core knowledge graph without duplicating data into RDF.

Overview

What this was about

Mimouni presented the Semantic Data Cloud, a containerised application that exposes distributed Darwin Core Data Package datasets through a single SPARQL endpoint without copying them into a triple store. Datasets are published as cloud-hosted Parquet files with an EML JSON-LD catalogue entry. The system selects relevant datasets by spatial, temporal, licence or maintenance filters, then rewrites SPARQL to SQL through an OWL 2 QL ontology with OBDA mappings based on the Darwin Core conceptual model, and executes it with DuckDB over HTTP. A live instance at data.qcbs.ca holds over 50 GBIF datasets and exposes an MCP server for LLM agents.

Why it matters. It offers a practical route to semantic, cross-dataset querying of the new Darwin Core Data Package without imposing a full RDF stack on data providers, and it is usable directly by LLM agents.

Key ideas

In the room

  • The ratified Darwin Core Data Package solves DwC-A's conflation of several entities in one table, but its semantics are implicit in table joins and it is hard to expose online.
  • The Darwin Core conceptual model defines relationships in natural language rather than in a formal ontology. An OWL ontology makes the semantics explicit and lets data be seen as a graph (e.g., a material entity evidencing an occurrence during an event).
  • Conventional RDF publication costs data duplication, fragile federated queries, a heavy RDF/ETL burden on providers, and schema drift.
  • Core principle: 'the data should be distributed, not the infrastructure'. Tables are exploded Parquet files on object storage instead of zipped CSVs.
  • An EML JSON-LD file per dataset carries asset hrefs pointing to the Parquet files (as in STAC catalogues) and supports dataset discovery with filters.
  • SPARQL is translated to SQL using an OWL 2 QL ontology and OBDA mappings, which guarantees rewritability while remaining a good fit for vocabularies like Darwin Core. DuckDB executes the SQL on remote Parquet over HTTP, so the virtual knowledge graph exists only for the query's duration.
  • The layers are separate: storage (Parquet), discovery (metadata catalogue), execution (DuckDB remote SQL), and planning/reasoning (SPARQL plus OWL).
Jump in

Notable moments

In their words

Transcript

Automatically generated captions can contain mistakes, especially in names and technical terms. Times are relative to the room recording.

Read the transcript ↓
Loading transcript…