Jul 2026· International Journal of Learner Corpus Research· 0 citations· 35 references
TL;DR
The efficiency metrics showed the tool to be effective in accurately identifying the target suffixes across proficiency levels, with a significant reduction in time and effort.
Abstract
This report presents the preliminary evaluation of
Morph
, a web-based tool for the automatic counting of 51 English noun derivational suffixes in a controlled corpus of Mexican learners of English with proficiency levels from A2 to C1 on the Common European Framework of Reference for Languages (CEFR) scale. The evaluation consisted of a quantitative analysis of agreement between human annotators and
Morph
, a qualitative analysis to obtain the sources of disagreement, and finally, the evaluation of the tool’s efficiency using precision, recall, and F1 score metrics. Results displayed a high level of agreement between the human annotators and
Morph
. The main sources of disagreement were found to be human errors, learner misspellings, and automatic tagger issues. Finally, the efficiency metrics showed the tool to be effective in accurately identifying the target suffixes across proficiency levels, with a significant reduction in time and effort.
We present CzeLeD, a large-scale lexical decision dataset for Czech containing 12,242 lemmas and inflected word forms derived from the HeCz self-paced reading corpus and supplemented with pseudowords. A total of 1,977 participants contributed over 849,000 trials, accompanied by demographic data and extensive linguistic annotation, including lemma and word-form frequencies, neighborhood density, phonotactic and orthotactic probability. Alongside the trial-level dataset, we provide lexical decision norms summarizing reaction times and accuracy using a transparent and reproducible trimming procedure. Data validation drew on both response accuracy and reaction times. Participants showed high overall accuracy (94.79%), and reaction times exhibited the expected strong inverse relationship with word form frequency. Importantly, CzeLeD can be directly linked to the HeCz corpus, enabling systematic comparison between word recognition in isolation and in sentential context. By combining broad lexical coverage, rich annotation, and seamless integration with an existing context-bound processing corpus, CzeLeD offers a powerful, reusable resource for investigating lexical processing, morphological complexity, and contextual effects in Czech and beyond.
J. Chromý, Markéta Ceháková, Mikuláš Preininger et al.· Scientific Data· 0 citations
This paper uses morphological feature annotations from the Universal Dependencies corpora to quantify information carried by morphology across 154 different datasets spanning 72 language varieties, and proposes an information-theoretic approach that measures how surprising morphological feature values are given a token’s part of speech or lemma in the corpus.
Hedvig Skirgård, S. Mann· Linguistics Vanguard· 0 citations
In this paper, we draw a comparison between linguists in training, a trained linguist, and annotations generated by large language models (LLMs) to find out if they struggle with complex linguistic phenomena in a similar way. For this purpose, we analyse evaluative language in spoken popular science discourse, with the example of a corpus of English TED talk transcripts. We focus on the Appraisal theory and its Attitude subsystem, including the categories (classes) of Affect, Judgement, and Appreciation. In this context, Appraisal theory is an example of a highly subjective annotation task, making it a suitable example for the study of complex annotation challenges. First, we assess human annotations on a sentence level in specific scientific domains. Then, we develop three prompts and compare them for model performance for the automatic classification of Appraisal classes. We assess the performance of three LLMs using the best-performing prompt and finetune the model, reaching an F1-score of 0.77. We find that models perform best compared to annotations conducted by the trained linguist, while linguists in training do not reach high agreement scores. We conclude that LLMs can aid in complex annotation task resolution, opening new pathways for the complex theories annotated and analyzed in digital humanities studies.
Mirela Imamović, Aenne Knierim, Khushi Pitroda et al.· 0 citations
Acquiring L2 prepositions is a complex process, as bilinguals often find it difficult to use the correct preposition in a target language such as English. While frequency of prepositions and prepositional phrases can help this acquisition process, additional factors such as L1 influence also play an important role. Furthermore, the acquisition of L2 prepositions follows a non-linear pattern, similar to U-shaped development in L2 acquisition, making investigations across different proficiency levels crucial in understanding the whole phenomenon. Using an error-coded learner corpus, this study investigated the frequencies of L2 preposition use and error types in the essays written by Turkish-English bilinguals. Three types of preposition errors (substitution, addition, and omission) were examined across different score bands corresponding to different writing proficiency levels. Overall results showed that although the frequency of prepositions accounted for their usage in L2 for the most common 10 prepositions, error types and how they differ across score bands are also essential in understanding the developmental pattern of L2 prepositions. In line with the literature, substitution and omission errors were the most common error types. Notably, omission errors demonstrated a statistically significant U-shaped developmental pattern, while substitution errors showed a similar but non-significant descriptive tendency across proficiency bands. Addition errors, on the other hand, decreased steadily with higher score bands. The current findings offer important insights into L2 preposition error types and development for fields like psycholinguistics, bilingualism, and corpus linguistics.
Unknown authors· Cankaya University journal o...· 0 citations
A narrative review compares contemporary Latin Natural Language Processing tools and their use cases for pedagogical applications and concludes that an integrated, multi-tool approach is most effective for supporting Latin pedagogy.
Aidan Han· American Journal of Student...· 0 citations
This work introduces PERCEPT, the first publicly available large-scale Persian-English code-mixed corpus annotated with Universal Dependencies POS tags for code-mixed words, and conducts the first comprehensive linguistic analysis of Persian-English code-mixing across multiple social media platforms.
Ghazal Kalhor, Zahra Jafari, Amirarsalan Shahbazi et al.· 0 citations