The Cross-Lingual Comprehension Gap (CLCG) is defined as the reduction in response quality when the same content and question are presented in a target language rather than in English.
Abstract
Language models are often evaluated as though capabilities demonstrated in English remain equally available when the same content is presented in other languages. Traditional multilingual benchmarks rarely isolate language while holding content, question, reference answer, model, and evaluation unit constant. We define the Cross-Lingual Comprehension Gap (CLCG) as the reduction in response quality when the same content and question are presented in a target language rather than in English. Using ParallelQA-18, a professionally human-translated parallel corpus, we evaluate five models from five laboratories on a stratified sample of 150 articles across 18 languages (English reference; Portuguese high-resource baseline; 16 targets spanning Joshi et al. 2020 classes 0-4). A within-item design varies only passage language. The primary estimator contrasts English versus pooled target-language Token-F1 micro-means on higher-complexity open-ended questions, with article-cluster bootstrap intervals. The primary pooled CLCG is 0.078 (95% CI 0.072-0.084), about a 17% reduction relative to the English score; the equal-language macro summary is 0.077. Net of Portuguese, the macro gap is 0.016 (95% CI 0.013-0.020). Language-level CLCG is negatively associated with Joshi resource class (rho = -0.594, p = 0.015, n = 16). In blinded paired human evaluations, higher-resource responses are preferred in 61.6% of decisive judgments (estimated preference probability 0.655, 95% CI 0.558-0.741). Capabilities shown in English should not be assumed to transfer equally to other languages; English-centered evaluations may overestimate quality for users of low-resource languages.
This work introduces M-GATE (Multilingual Grammar, Accuracy in Translation, and Efficiency), a benchmark of linguistic proficiency spanning 30 typologically diverse languages from high- to low-resource, and evaluates over 50 models in more than 80 configurations.
Tomáš Burkert, Angelika Peljak-Łapińska, David Zelený· 0 citations
J-PragEval-v0 is introduced, a minimal-pair benchmark isolating four such phenomena from surface fluency, and Pragmatic Representation Steering is specified, a parameter-free inference-time method that edits residual-stream activations along the class-mean-difference directions probing identifies.
In every setting, pre-adaptation on related auxiliary languages yields no practically meaningful improvements once as little as one hour of target-language data is available, suggesting that relatedness alone may not reliably predict transfer gains in large multilingual ASR, or constitute an effective strategy for extending such models to low-resource languages.
A. Florian, C. Amol, Hope Kerubo Ombaba et al.· 0 citations
Large language models access knowledge inconsistently across languages, but to what extent do they differ in their skill sets when interacting with different languages? This work quantifies cross-lingual skill inconsistency orthogonally from knowledge and general benchmark performance. We do this via multilingual self-play: two instances of the same model compete in a text-based game, each interacting through a different language interface. Since the model, opponent, rules, state space, and available actions remain fixed, this setting isolates the effect of language on the model's realized behavior. We build a multilingual extension to TextArena and evaluate three open-weight models across eight languages and six games covering spatial reasoning, imperfect information, resource allocation, and repeated interaction. We find that the same model can exhibit markedly different playing strength across languages, with systematic variation in win--loss margins, invalid actions, and strategic tendencies. Detailed analyses reveal language-specific failures in spatial reasoning, card-conditioned decisions, and optimal move selection. In some settings, changing only the intermediate reasoning language recovers much of the lost performance, suggesting that language can affect different stages of the decision process. These results show that skill discrepancies are a measurable major roadblock in the development of truly multilingual models. Better understanding these discrepancies can help us design models that perform more equitably across languages.
Bobby Cheng, Adam Gaber, Zhengzhe Liu et al.· 0 citations
This paper examines the types of English-language questions posed by learners on HiNative, a global peer-to-peer language learning platform. While online question-and-answer platforms have been widely adopted for language learning, the linguistic focus and distribution of learner-generated questions in self-access digital environments remain underexplored. Drawing on a linguistic content analysis approach informed by learner autonomy, the study systematically classifies learners’ questions across key linguistic domains. Using qualitative content analysis, a dataset of 783 learner-generated inquiries was categorized into key linguistic domains (grammar, vocabulary, usage, pronunciation, idioms, translation, comparison, formality/contextual use, and other). The findings reveal a predominant focus on meaning-oriented questions, particularly translation (34.2% of all questions) and vocabulary (21.5%), while grammar and pronunciation queries were comparatively rare. This distribution suggests that learners prioritize communicative meaning, lexical expansion, and pragmatic appropriateness over formal rule-focused learning. Such inquiry patterns highlight learners’ reliance on their first language as a bridge to the target language and their pursuit of precise and natural usage in authentic contexts. These results contribute to second language acquisition research by offering insights into learner priorities in informal, technology-mediated settings, with implications for the design of adaptive learning resources and learner support systems.
R. Alsharif· Theory and Practice in Langu...· 0 citations
While large language models (LLMs) are increasingly used across social domains, current bias evaluations often rely on demographic proxies such as names, pronouns, and social categories. Linguistic variety itself receives less attention as an evaluative variable. This study therefore uses sociolinguistic variety to examine whether three LLMs respond differently to semantically equivalent prompts in General American English and Irish English. Thirty prompt pairs were administered twice to GPT-5, Claude, and Gemini, yielding 360 responses. Across the 180 paired comparisons, 92.8% differed in word count, although the direction varied by model: Claude produced longer Irish English responses on average, whereas GPT-5 and Gemini produced shorter responses. Claude explicitly referenced Irish English features in 55.0% of its Irish English responses, compared with 0.0% for GPT-5 and 1.7% for Gemini. No explicit correction, refusal, or researcher-observed tone shift occurred in either condition. These findings show variety-conditioned differences in response behavior, but they do not by themselves establish discriminatory biasThe study supports linguistic variety as an additional dimension for LLM differential-behavior evaluation and identifies model- and feature-specific patterns that warrant further evaluation with independent raters and additional varieties.
Awad H. Alshehri, N. Jaballah· Journal of Intelligent Decis...· 0 citations