Skip to content

PolyglotQL: A Pipeline for Multilingual Text-to-SPARQL Dataset Generation

2026 · International Conference on Language Resources and Evaluation · pp. 6674-6684 · 0 citations · 26 references
Computer Science

TL;DR

PolyglotQL provides an extensible and modular architecture that aggregates, normalizes, and augments heterogeneous question–SPARQL pairs from established text-to-SPARQL datasets, highlighting the benefits of structured context in multilingual semantic parsing.

View source

Similar papers

Open access 2026

TextLens & LeTTuce: Automated Corpus Annotation and Multilingual Tagging as a Service

We present TextLens , a web-based platform for automated linguistic annotation designed to lower technical barriers for researchers in digital humanities, linguistics and translation studies. Hosted by the Dutch Language Institute (INT), TextLens allows users to upload and annotate corpora in a variety of formats (.txt...

Cynthia Van Hee, Jonas Doumen, Vincent Prins et al. · 0 citations
#artificial intelligence Preprint Sep 2026

Build2SPARQL: A Large-Scale Text-to-SPARQL Benchmark Dataset for Building Knowledge Graph Querying

Building automation systems are increasingly represented as semantic knowledge graphs (KGs) using ontologies such as Brick and ASHRAE 223P, creating a machine-readable substrate for artificial-intelligence applications. One promising application is translating natural-language questions into SPARQL (text-to-SPARQL), wh...

Wooyoung Jung · 0 citations
Open access 2026

PARSEME 2.0 Multilingual Corpus of Multiword Expressions

We present edition 2.0 of the PARSEME multilingual corpus annotated for multiword expressions (MWEs), resulting from efforts of the PARSEME community towards universality-driven modeling of idiomaticity. With respect to previous editions, we extend the annotation scope to all syntactic MWE categories: verbal, nominal,...

Agata Savary, Manon Scholivet, Carlos Ramisch et al. · 1 citation
Preprint Aug 2026

L3Cube-IndicQuest v2: A Large-Scale Multilingual Benchmark for Evaluating Factual Knowledge of Large Language Models Across Indic Languages

We present L3Cube-IndicQuest v2, a large-scale gold-standard multilingual question-answering benchmark for evaluating the India-specific factual knowledge of Large Language Models (LLMs). The benchmark comprises 3,471 curriculum-grounded English question--answer pairs spanning nine domains, curated from educational cur...

Rinit Jain, Tirthraj Mahajan, Advait Joshi et al. · 0 citations
#machine learning Preprint Oct 2026

FALCON: A Model and Dataset Agnostic Framework for Synthetic Data Generation for NL2SQL Pairs

Relational databases are among the most widely deployed forms of structured knowledge, and natural language access to them requires grounding language onto schema entities and relations while handling the ambiguity inherent in how people phrase requests. Existing synthetic NL-to-SQL data generation methods largely igno...

Darian Lee, Shannon Rumsey, Jack St. Clair et al. · 0 citations
Preprint Aug 2026

MGAL: A Multilingual Granularity-Aware Long-Context Benchmark

MGAL is the first multilingual, granularity- and position-aware long-context benchmark, constructed from United Nations reports spanning 8K to 128K tokens across the six official UN languages, and finds that LLMs perform well at word-level tasks but struggle with coarser-grained ones.

Chunhan Li, Chenglin Xu, Zongyang Zhang et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.