Skip to content
Open access

Morphologically Annotated Lexical Decision Data for 12,242 Czech Word Forms

Jul 2026 · Scientific Data · 0 citations

Abstract

We present CzeLeD, a large-scale lexical decision dataset for Czech containing 12,242 lemmas and inflected word forms derived from the HeCz self-paced reading corpus and supplemented with pseudowords. A total of 1,977 participants contributed over 849,000 trials, accompanied by demographic data and extensive linguistic annotation, including lemma and word-form frequencies, neighborhood density, phonotactic and orthotactic probability. Alongside the trial-level dataset, we provide lexical decision norms summarizing reaction times and accuracy using a transparent and reproducible trimming procedure. Data validation drew on both response accuracy and reaction times. Participants showed high overall accuracy (94.79%), and reaction times exhibited the expected strong inverse relationship with word form frequency. Importantly, CzeLeD can be directly linked to the HeCz corpus, enabling systematic comparison between word recognition in isolation and in sentential context. By combining broad lexical coverage, rich annotation, and seamless integration with an existing context-bound processing corpus, CzeLeD offers a powerful, reusable resource for investigating lexical processing, morphological complexity, and contextual effects in Czech and beyond.

Read PDF

Similar papers

Open access Jul 2026

How novel are low-frequency words?

This study explores the relationship between low-frequency lexical forms and lexical innovation by examining infrequent and non-lexicalised adjectives formed with the suffix -able (e.g., jokeable, trickable). Employing a corpusbased analysis of the 20-billion-word News on the Web (NOW) corpus (2010–2025), we identified over 800 infrequent Xable adjectives. After filtering for orthographical errors, brand names, and incorrect forms, a dataset of novel lexical items was isolated using the OED for evidence of attestation. Morphological patterns and usage contexts were analysed, highlighting factors such as (1) frequency distribution over 2010–2025, (2) geographical coverage, (3) contextual anchoring in the collocate environment, and 4) base form and derivative co-occurrence. Findings contribute insights into lexical innovation, morphological productivity, and the stages of lexical integration from a dynamic usage-based perspective (Schmid 2020).

Chris A. Smith · 0 citations
Open access Aug 2026

An improved metric for estimating morphological information in corpora

This paper uses morphological feature annotations from the Universal Dependencies corpora to quantify information carried by morphology across 154 different datasets spanning 72 language varieties, and proposes an information-theoretic approach that measures how surprising morphological feature values are given a token’s part of speech or lemma in the corpus.

Hedvig Skirgård, S. Mann · 0 citations
Open access Jul 2026

Low resource word sense disambiguation in Oromo with fine tuned small transformers.

Results show that contextual transformer representations are quite successful for low-resource WSD, although there is still a significant class imbalance that limits performance.

Liyachew Edeti, Million Meshesha, Feda Negesse · 0 citations
Jul 2026

The Arabic Lexicon Project: Lexical decision data for 10,000 Modern Standard Arabic words

The current study developed, validated, and illustrated the advantage of a lexical decision database for MSA in three (virtual) experiments, and showed that the database largely replicates well-documented effects in visual word recognition (lexicality, word length, word frequency, word neighborhood size).

Alaa Alzahrani, Hassan Alshumrani, Wafa Aljuaythin et al. · 0 citations
Open access Jul 2026

Nominalization Patterns in Tombatu Language: A Generative Morphological Analysis of Affixal Noun Formation

Regional languages preserve complex grammatical systems that reveal how communities organize experience, identity, and cultural knowledge. However, nominalization in Tombatu, an underdescribed Austronesian language of Minahasa, has not been systematically examined through an explicit rule-based morphological framework. This study addresses this gap by investigating the structural patterns, semantic functions, and generative mechanisms of affixal noun formation in Tombatu. Employing a descriptive-taxonomic design, data were collected over six months in Tombatu Village from proficient speakers through naturally occurring speech, elicitation, recording, note-taking, and informant validation. The final dataset found unique nominalized forms. These forms were classified into three major processes comprising 11 noun-forming prefixes, one noun-forming suffix, and 14 affix combinations, yielding 26 identified affixal configurations. The analysis applied Word Combination Rules, Derivation Rules, the Boundary Insertion Convention, and relevant morphological constraints. The findings show that prefixation exhibits the greatest structural and semantic variation in the dataset, producing agentive, person-denoting, instrumental, object-related, and result-related nouns. For example, /tapə-/ combines with /lukuʔ/ ‘to drink’ to form /tapəlukuʔ/ [tapəlukuʔ] ‘drinker’. Suffixation is more restricted, with /-an/ forming locative nouns, as in /tawoi/ ‘to work’ plus /-an/, which produces /tawoian/ [tawoian] ‘workplace’. Complex affix combinations form nouns denoting places, processes, objects, states, results, and entities. This study provides the first systematic generative account of Tombatu nominalization, expands empirical evidence for Austronesian word formation, strengthens regional-language documentation, and offers a linguistic basis for language maintenance and potential contrastive morphological activities in multilingual education.

Verra E. Manangkot, J. Liuw, L. Amalputra et al. · 0 citations