This work strengthens an already valuable community resource by aligning its artifacts with the RDF standards, broadening the set of engines that can be fairly and reproducibly compared, and raises the question of how the results of federated queries under automatic source selection can be made reproducible.
Abstract
LargeRDFBench is one of the most comprehensive benchmarks for evaluating federated SPARQL query engines, combining real, interlinked datasets with a rich query suite that has made it a reference point for the community. Evaluations of federated engines are published by comparing engine results against the benchmark's expected results, so those expected results must themselves be reproducible. Moreover, several of its data dumps violate the RDF specifications, so only engines that parse RDF leniently can host them, and its expected results are distributed in an ad hoc format. We identify, categorize and repair these data-quality issues with a reproducible cleaning pipeline, producing standards-conformant serializations of every affected dataset. Furthermore, we re-encode the benchmark's expected results in the W3C SPARQL 1.1 Query Results JSON Format and correct their discrepancies. Every dataset now parses under strict RDF parsers, and the expected results are machine-verifiable through a standard format, extending the benchmark's reach to the full range of conformant engines while staying faithful to the original data. Reproducing the expected results end-to-end with an independent implementation uncovers corruption in the published reference, and discrepancies between our results and the original ones, some not trivial to resolve, others open questions. We further perform a preliminary comparison, not previously explored, of ASK- and COUNT-based source selection in the FedX algorithm. This work strengthens an already valuable community resource by aligning its artifacts with the RDF standards, broadening the set of engines that can be fairly and reproducibly compared. We also raise the question of how the results of federated queries under automatic source selection can be made reproducible.
This work proposes RENSA, a federated SPARQL query generation framework that leverages an extension of SPARQL Builder Metadata (SBM), and demonstrates that RENSA infers class and authority constraints for query variables, enabling the identification of data sources even across heterogeneous endpoints.
Victor Eiti Yamamoto, Hideaki Takeda, Yasunori Yamamoto· 0 citations
MGQL is presented, the first mechanized, small-step operational semantics for a substantial read-only fragment of GQL that is grounded in the ISO/IEC 39075 standard, and it is proved that the type system is sound, ensuring an end-to-end guarantee of well-formed queries yielding results that conform to their declared sc...
Aditya Thimmaiah, Tong-Tong Lin, Milos Gligoric· Proceedings of the ACM on Pr...· 1 citation
This work demonstrates how validation can serve not only as a quality assurance mechanism, but as a central organizing principle for workflow automation, enabling scalable and reliable content management across a distributed XML ecosystem.
Kin Ng, Lisandro Gonzalez, Stacy Lathrop· Balisage Series on Markup Te...· 0 citations
The results show that structured RDF pipelines currently provide the most stable integration behavior, whereas JSON and text pipelines remain more sensitive to errors in mapping, extraction, and linking.
OntoExpand is introduced, a new methodology for ontology expansion that uses SPARQL CONSTRUCT queries as an efficient alternative to conventional reasoning techniques that improves performance, reduces computational overhead, is pattern driven allowing a more granular expansion control and seamlessly integrates with SP...
Vitor Lelis, N. Leite, José Carlos et al.· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.