Skip to content
Preprint

Strengthening LargeRDFBench for Interoperable Federated SPARQL Evaluation

Aug 2026 · 0 citations · 25 references
Computer Science

TL;DR

This work strengthens an already valuable community resource by aligning its artifacts with the RDF standards, broadening the set of engines that can be fairly and reproducibly compared, and raises the question of how the results of federated queries under automatic source selection can be made reproducible.

Abstract

LargeRDFBench is one of the most comprehensive benchmarks for evaluating federated SPARQL query engines, combining real, interlinked datasets with a rich query suite that has made it a reference point for the community. Evaluations of federated engines are published by comparing engine results against the benchmark's expected results, so those expected results must themselves be reproducible. Moreover, several of its data dumps violate the RDF specifications, so only engines that parse RDF leniently can host them, and its expected results are distributed in an ad hoc format. We identify, categorize and repair these data-quality issues with a reproducible cleaning pipeline, producing standards-conformant serializations of every affected dataset. Furthermore, we re-encode the benchmark's expected results in the W3C SPARQL 1.1 Query Results JSON Format and correct their discrepancies. Every dataset now parses under strict RDF parsers, and the expected results are machine-verifiable through a standard format, extending the benchmark's reach to the full range of conformant engines while staying faithful to the original data. Reproducing the expected results end-to-end with an independent implementation uncovers corruption in the published reference, and discrepancies between our results and the original ones, some not trivial to resolve, others open questions. We further perform a preliminary comparison, not previously explored, of ASK- and COUNT-based source selection in the FedX algorithm. This work strengthens an already valuable community resource by aligning its artifacts with the RDF standards, broadening the set of engines that can be fairly and reproducibly compared. We also raise the question of how the results of federated queries under automatic source selection can be made reproducible.

View source

Similar papers

#natural language process... Preprint Aug 2026

RENSA: Rich Environment Metadata to Navigate Shared and Distributed Endpoints for Automated Federated SPARQL Query Generation

This work proposes RENSA, a federated SPARQL query generation framework that leverages an extension of SPARQL Builder Metadata (SBM), and demonstrates that RENSA infers class and authority constraints for query variables, enabling the identification of data sources even across heterogeneous endpoints.

Victor Eiti Yamamoto, Hideaki Takeda, Yasunori Yamamoto · 0 citations
#small language model Open access Aug 2026

MGQL: An Executable, Small-Step Semantics of GQL

MGQL is presented, the first mechanized, small-step operational semantics for a substantial read-only fragment of GQL that is grounded in the ISO/IEC 39075 standard, and it is proved that the type system is sound, ensuring an end-to-end guarantee of well-formed queries yielding results that conform to their declared sc...

Aditya Thimmaiah, Tong-Tong Lin, Milos Gligoric · 1 citation
Review

Validation-Driven Automation in a Federated XML Ingest Pipeline: The NCBI Bookshelf Case

This work demonstrates how validation can serve not only as a quality assurance mechanism, but as a central organizing principle for workflow automation, enabling scalable and reliable content management across a distributed XML ecosystem.

Kin Ng, Lisandro Gonzalez, Stacy Lathrop · 0 citations

KGpipe: Generation of Pipelines for Data Integration into Knowledge Graphs

The results show that structured RDF pipelines currently provide the most stable integration behavior, whereas JSON and text pipelines remain more sensitive to errors in mapping, extraction, and linking.

Marvin Hofer, Erhard Rahm · 1 citation

OntoExpand: SPARQL-Based Ontology Expansion and Reasoning

OntoExpand is introduced, a new methodology for ontology expansion that uses SPARQL CONSTRUCT queries as an efficient alternative to conventional reasoning techniques that improves performance, reduces computational overhead, is pattern driven allowing a more granular expansion control and seamlessly integrates with SP...

Vitor Lelis, N. Leite, José Carlos et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.