Skip to content
Open access

Corpus-Based Identification and Profiling of Vocabulary in Linguaskill Speaking Tests: A Methodological Framework for Adaptive Test Validation

2026 · International journal of research and innovation in social science · Vol 10, pp. 17207-17224 · 0 citations

TL;DR

A complete corpus-based methodological framework for identifying, classifying, and profiling the vocabulary of the Linguaskill Speaking Test and contributes a replicable protocol for lexical validation of adaptive speaking tests and a benchmark synthesis for interpreting its future results.

Abstract

Computer-adaptive tests such as Cambridge Linguaskill adjust item difficulty dynamically to match test-taker proficiency, and their speaking components are now used for high-stakes admission and exit decisions, including in Malaysian higher education. Yet little research has examined the vocabulary load and lexical demands of adaptive speaking assessment, leaving open whether the vocabulary elicited at each reported level is congruent with the Common European Framework of Reference (CEFR) levels the scores claim to represent. This article develops a complete corpus-based methodological framework for identifying, classifying, and profiling the vocabulary of the Linguaskill Speaking Test. The framework specifies the compilation of a specialised corpus of transcribed candidate responses stratified across CEFR levels, supplemented by official test preparation materials, and a three-lens profiling procedure using LexTutor for BNC-COCA word-family analysis, the New General Service List for high-frequency lemma coverage, and Text Inspector for CEFR-aligned profiling against the English Vocabulary Profile. Operational criteria are defined for distinguishing core from peripheral vocabulary at each adaptive level, and decision rules are specified for classifying each level as aligned, over-demanding, or under-demanding relative to CEFR expectations, benchmarked against published lexical coverage thresholds for spoken English. The central hypothesis, motivated by prior coverage research, is that the lexical demands elicited by adaptive speaking tasks will not rise in step with claimed CEFR levels. The framework is grounded in an argument-based approach to validation and articulates the washback, fairness, and pedagogical stakes of the alignment question. The article contributes a replicable protocol for lexical validation of adaptive speaking tests and a benchmark synthesis for interpreting its future results.

Read PDF

Similar papers

Open access Jul 2026

Constructing the LERWL: A Corpus-Based Approach to Identifying Technical Vocabulary in Language Education Research

Technical vocabulary plays a crucial role in ESP learners’ lexical development once general service and academic vocabulary has been established. The present study constructed the Language Education Research Corpus (LERC), an 8,646,901-word corpus compiled from the top 10 Quartile 1 (Q1) Scopus-indexed journals in Language Education. On the basis of this corpus, the Language Education Research Word List (LERWL) was developed through four systematic procedures. High-frequency items were first identified, after which a range criterion was applied to retain words occurring in at least 50% of the target journals. Lexical profiling was subsequently conducted to isolate discipline-specific vocabulary, excluding items listed in the GSL and the AWL. In addition to these quantitative procedures, expert judgment was employed as a qualitative validation measure to ensure disciplinary relevance. The resulting LERWL was then evaluated against (a) the LERC, achieving 5.28% coverage, and (b) an independent corpus of approximately one million words from the same field, achieving 4.75% coverage. These findings indicate that the LERWL provides a substantial and pedagogically meaningful level of lexical coverage within the discipline.

Supakorn Phoocharoensil · 0 citations
Aug 2026

Analysis of the Efficiency of Subword Tokenizers in a Low-Resource Linguistic Environment: Implementation Experience for the Tajik Language

The experimental results revealed the strengths and weaknesses of various approaches to subword segmentation and identified the most effective tokenization strategies under the conditions of the morphological complexity of the Tajik language.

M. Arabov, S. Khaybullina · 0 citations
Open access Jul 2026

Lexical diversity and CEFR vocabulary in critical academic reading passages: A conceptual replication

Background and Purpose: Critical academic reading skills demand learners to engage in higher-order reading skills which require advanced lexical knowledge. The two dimensions of lexical knowledge, namely lexical diversity and vocabulary variety, however, are mostly assessed in isolation using unstable measurement instruments, hence giving misleading results. The purpose of this study was to examine these two dimensions in critical academic reading passages through an integrated approach using a more reliable measurement instrument. Methodology: The corpus consisted of 32 selected reading passages. They were first cleaned to remove any interference, such as numbering that might affect the lexical calculation, before being converted to plain text format. The Text Inspector was then used to generate the Measure of Textual Lexical Diversity (MTLD) and CEFR vocabulary difficulty levels for all passages. The scores were then analysed descriptively to examine the patterns of lexical diversity and CEFR vocabulary distributions across the passages. Additionally, Spearman’s rho correlational analysis was conducted to examine the relationship between MTLD scores and CEFR-based vocabulary levels. Findings: All 32 examined passages generally exhibit high lexical diversity, with many common and basic words. Lexical diversity was found to increase when advanced vocabulary was added to the passage, indicating a significant positive correlation between the two dimensions. Contributions: The integrated approach in lexical analysis adopted in this study results in a comprehensive picture of lexical complexity rather than using the single metric alone, hence offering a clearer and fairer approach for instructors in evaluating and selecting reading passages to be used with students. Keywords: CEFR profiling, critical academic reading, lexical difficulty, lexical diversity, Measure of Textual Lexical Diversity (MTLD). Cite as: Aziz, A., Syed Ahmad, T. S. A., Awang, S., Abdul Aziz, R., Azlan, N. A., & Ahmad, S. N. (2026). Lexical diversity and CEFR vocabulary in critical academic reading passages: A conceptual replication. Journal of Nusantara Studies, 11(2), 228-242. https://dx.doi.org/10.24200/jonus.vol11iss2pp228-242

Anealka Aziz, Tuan Sarifah Aini Syed Ahmad, Suryani Awang et al. · 0 citations
Preprint Aug 2026

Measuring the Tokenization Premium: A Cost Audit for Underserved Language Communities

Large language models are increasingly deployed as general-purpose educational and technical assistance systems, but their underlying infrastructure does not treat languages equally. One underexamined source of disparity is tokenization: semantically equivalent content can require substantially different token counts across languages, affecting API cost, latency, and usable context length before a model is invoked. We introduce the Tokenization Equity Audit (TEA), a reproducible benchmark for measuring tokenization premiums in technical tutoring content. TEA evaluates three widely used tokenizers, GPT-4o's o200k base, Qwen2.5-7B, and Mistral-7B, on a 120-item Python debugging corpus translated from English into Bengali, Hindi, Arabic, Tamil, and Yoruba. Bengali and Hindi serve as the primary validated cases, while the remaining languages provide exploratory cross-script and cross-family comparisons. Across this corpus, Bengali requires (1.56\times) as many GPT-4o tokens as English, reducing a nominal 128k-token context window to an effective 82k-token English-equivalent capacity for the same semantic content. With the Qwen2.5 and Mistral tokenizers, Bengali requires up to (4.5\times) the English token count. Yoruba, despite using the Latin script, exhibits the highest GPT-4o tokenization premium at (2.37\times), indicating that tokenization inequity cannot be explained by script family alone. These results demonstrate that tokenization can create measurable economic and functional barriers, highlighting the need to treat tokenization as an equity-relevant infrastructure layer for underserved language communities, particularly where educational systems depend on low-cost or offline-capable AI tools.

Avijit Roy, Proma Roy, Hrishitva Patel · 0 citations
Jul 2026

The KSAUHS learner corpus

The KSAU-HS Learner Corpus is a longitudinal corpus of EFL tertiary writing that complies with the FAIR principles. Collection began in 2022 and captures writing development during a period of emerging language technologies (2022–24). The corpus contains over 856,907 tokens across 2,387 texts produced by 157 preparatory year university students, within a CEFR-aligned program with an instructional range of approximately A2-B2. Texts span four trimesters and include rhetorical modes such as cause-and-effect, argumentation, and summarisation. Metadata includes years of English schooling, other languages spoken, and preferred reference tools. The corpus enables research into writing development to inform EAP pedagogy and assessment. Avenues for investigation include lexico-grammatical development, cross-linguistic influence, individual differences, and the impact of task conditions and language technologies. This resource promises data-driven insights into the textual features and factors that characterise EFL writing proficiency, and plans are in place to expand its size, representation, and accessibility.

Eman Al Nafjan, Alaa Alfelaij, N. Alfawaz et al. · 0 citations