Jul 2026· International Journal of Learner Corpus Research· 0 citations· 21 references
Abstract
The KSAU-HS Learner Corpus is a longitudinal corpus of EFL tertiary writing that complies with the FAIR
principles. Collection began in 2022 and captures writing development during a period of emerging language technologies (2022–24).
The corpus contains over 856,907 tokens across 2,387 texts produced by 157 preparatory year university students, within a
CEFR-aligned program with an instructional range of approximately A2-B2. Texts span four trimesters and include rhetorical modes
such as cause-and-effect, argumentation, and summarisation. Metadata includes years of English schooling, other languages spoken,
and preferred reference tools. The corpus enables research into writing development to inform EAP pedagogy and assessment. Avenues
for investigation include lexico-grammatical development, cross-linguistic influence, individual differences, and the impact of
task conditions and language technologies. This resource promises data-driven insights into the textual features and factors that
characterise EFL writing proficiency, and plans are in place to expand its size, representation, and accessibility.
This empirical corpus-based TESOL study examines mode and proficiency effects in EFL learner production through a planned comparison of written essays and spoken monologues in the International Corpus Network of Asian Learners of English (ICNALE). The study focuses on how lexical, grammatical, and selected discourse features vary across written and spoken learner output, and how these patterns interact with proficiency level. Using a quantitative comparative corpus design, the study will analyze measures such as word count, lexical density, lexical diversity, academic vocabulary use, syntactic complexity, discourse markers, stance markers, and selected fluency-related transcript features where available. At this draft stage, the final ICNALE sample, feature counts, statistical outputs, and figures have not yet been supplied; therefore, all numerical findings are marked for author completion and verification. The manuscript is structured to allow the author to insert verified values after registering for ICNALE access, selecting comparable prompts, cleaning the dataset, and completing the analysis. The expected contribution is a mode-sensitive account of learner production that may inform integrated TESOL writing and speaking instruction, proficiency-sensitive assessment, and corpus-informed materials development. The paper differs from corpus-literacy or teacher-education review work by focusing on empirical learner-language evidence rather than teachers’ corpus-use knowledge.
Raga Fatthallah· Al-Farooq Journal of Science...· 0 citations
The new version retains the core concordancing, n-gram and disciplinary variation functions of its predecessor while introducing a range of significant enhancements including a rebuilt Python/FastAPI backend, five collocation association measures, and – most substantially – a fully integrated large language model (LLM) assistant that reads users’ actual search results to provide evidence-grounded linguistic commentary.
This study investigates how French secondary school learners of English
as a Foreign Language (EFL)(14 year-olds, levels A1+ to B1+) develop the ability to refer to
the past during social interactions. Based on a corpus of video-recorded peer
interactions, our cross-sectional study integrates tools from an enunciative theory
(theory of predicative and enunciation operations) and conversation analysis for second
language acquisition. We identified key stages in the acquisition of past tense forms,
from reliance on adverbials and verb stems to the use of more complex past verb forms. Our
findings align with Klein and Perdue’s Basic Variety concept (
1997
), suggesting a ‘natural’ interlanguage development, even in
tutored settings where a different linguistic progression is implemented in the programme
of instruction. By valuing learner productions and focusing on interactional dynamics, our
research highlights the natural progression in EFL learners’ use of various forms when
referring to the past, contributing to a deeper understanding of their linguistic
development and providing empirical support for the view that grammar emerges from social
interaction.
Pascale Manoïlov, Agnès Leroux· Language, Interaction and Ac...· 0 citations
It is well-established that second/foreign language learning requires exposure to meaningful input. Thanks to advancing technologies, one way to achieve this in language teaching contexts is via online corpus tools. Given the potential pedagogical gains these tools offer for language skills (e.g., vocabulary, grammar, pronunciation, and writing), this study aims to synthesise the findings of research on five popular corpus tools – Sketch Engine, SkELL, PlayPhraseMe, Fraze.it, and CorpusMate. Following the PRISMA guidelines for systematic reviews, this study included five research papers from the Web of Science and Scopus databases until 2024. Applying qualitative and quantitative content analyses, the study found that corpus consultation promotes language skills and learner autonomy, supports contextually appropriate language use, and improves sensitivity to authentic language structures. Despite the promising outcomes, the scant number of studies analysed with some methodological issues and contextual constraints undermines the generalisability of the results. The study underscores the urgent need for more longitudinal, comparative, and large-scale research on using these corpus tools to harness them fully in language education. It also proposes that when systematically implemented, corpus tools may serve as powerful resources for language learning and teaching.
I. Topal· Journal of language research· 0 citations
The present study introduces an open-access corpus of Korean second language argumentative writing (L2K-ARG)
produced by university-level learners with different first-language backgrounds. The corpus contains 484 essays written in
response to three standardized prompts under timed, prompt-controlled settings. Its design was informed by the framework proposed
by
Egbert et al. (2022)
, which defines representativeness based on how well a corpus
reflects the target domain and the distribution of linguistic features within that domain. To characterize the learners presented
in the corpus, the dataset includes general Korean proficiency as well as writing proficiency scores based on an analytic rubric.
This resource aims to complement existing Korean learner corpora by providing a publicly available, prompt-controlled
argumentative writing samples with learner proficiency information, supporting reproducible research in L2 Korean.
Hakyung Sung, Gyu-Ho Shin, Boo Kyung Jung et al.· International Journal of Lea...· 0 citations