Jul 2026· Arbeiten aus Anglistik und Amerikanistik / Agenda: Advancing Anglophone Studies· 0 citations
Abstract
This study explores the relationship between low-frequency lexical forms and lexical innovation by examining infrequent and non-lexicalised adjectives formed with the suffix -able (e.g., jokeable, trickable). Employing a corpusbased analysis of the 20-billion-word News on the Web (NOW) corpus (2010–2025), we identified over 800 infrequent Xable adjectives. After filtering for orthographical errors, brand names, and incorrect forms, a dataset of novel lexical items was isolated using the OED for evidence of attestation. Morphological patterns and usage contexts were analysed, highlighting factors such as (1) frequency distribution over 2010–2025, (2) geographical coverage, (3) contextual anchoring in the collocate environment, and 4) base form and derivative co-occurrence. Findings contribute insights into lexical innovation, morphological productivity, and the stages of lexical integration from a dynamic usage-based perspective (Schmid 2020).
We present CzeLeD, a large-scale lexical decision dataset for Czech containing 12,242 lemmas and inflected word forms derived from the HeCz self-paced reading corpus and supplemented with pseudowords. A total of 1,977 participants contributed over 849,000 trials, accompanied by demographic data and extensive linguistic annotation, including lemma and word-form frequencies, neighborhood density, phonotactic and orthotactic probability. Alongside the trial-level dataset, we provide lexical decision norms summarizing reaction times and accuracy using a transparent and reproducible trimming procedure. Data validation drew on both response accuracy and reaction times. Participants showed high overall accuracy (94.79%), and reaction times exhibited the expected strong inverse relationship with word form frequency. Importantly, CzeLeD can be directly linked to the HeCz corpus, enabling systematic comparison between word recognition in isolation and in sentential context. By combining broad lexical coverage, rich annotation, and seamless integration with an existing context-bound processing corpus, CzeLeD offers a powerful, reusable resource for investigating lexical processing, morphological complexity, and contextual effects in Czech and beyond.
J. Chromý, Markéta Ceháková, Mikuláš Preininger et al.· Scientific Data· 0 citations
Conversion is a fertile resource of word-formation in English. Its degree of productivity is often reflected in the high incidence of converted lexemes in language resources, for instance dictionaries and corpora. With the main aim of surveying the alleged relationship between morphological productivity and frequency, this study examines English verbs derived by denominal conversion between 1990 and 2019. The analysis of a dataset compiled from the Oxford English Dictionary (OED) and the Corpus of Contemporary American English (COCA) reveals a concentration of converted verbs within the lowest frequency bands. The scarcity of converted verbs in mid-to-high frequency bands suggests that, while morphologically viable, they rarely achieve broad lexical entrenchment. This underscores the often ephemeral nature of the output of highly productive word-formation processes, with many formations remaining marginal or specialized in use, contingent upon their semantic categorization. It is noteworthy that, while the OED documents a significant number of converted verbs, the COCA data indicates that many of these remain peripheral in discourse. The prominence of converted verbs in the OED’s lowest bands suggests their recognition without widespread usage, consistent with prior observations that conversion frequently yields nonce formations and contextdependent lexical items. The semantic distribution of converted verbs further elucidates their productivity and frequency patterns. Instrumental verbs dominate, reflecting technological influence, while performative verbs also show remarkable representation. Overall, the findings confirm that denominal conversion is a productive but predominantly low-frequency process, shaped by morphosyntactic and semantic constraints.
Jesús Fernández-Domínguez· Arbeiten aus Anglistik und A...· 0 citations
This study examines the four major word classes (viz. noun, verb, adjective and adverb) in Lohorung, a Kirat Rai (Tibeto-Burman) language of north-eastern Nepal, based on their semantic, syntactic, and morphological properties, based on Givón’s formal and functional approach. Data were collected through elicitation with native speakers and textual analyses and were analysed with reference to the author’s native-speaker intuition. The findings indicate that Lohorung exhibits prototypical nouns with their multiple features and non-prototypical nouns as well. Morphologically, nouns are marked by the specific affixes, <-ʈsi> encodes dual and <-i> encodes plural. Lohorung employs three types of classifiers, namely <-tsi> <-kɔ>, and <-pɑŋ/pɑ>. Verbs are clause-final and serve as main predicates, typical of Tibeto-Burman languages. Lohorung has three numbers and persons systems with clusivity. Manner adverbials may take , and some adverbs modify adjectives. This study concludes that the analysis of the four major word classes enhances the grammatical description of the Lohorung language and contributes to functional-typological research on lexical classification.
Diwas Rai· Journal of Indigenous Knowle...· 0 citations
This study examines the performance of the state-of-the-art MARBERT model in identifying the lexical/pragmatic category associated with emoji use on X within a digital pragmatics approach (DPA). A net corpus of 15856 Colloquial Arabic (CA) posts containing emojis was collected from X using Python. The texts were tokenized and normalized into 4 lexical categories, namely noun_norm, verb_norm, adj_norm, and adverb_norm, and 2 pragmatic/structural categories, question_norm and exclamation_norm. MARBERT was finetuned and optimized to identify which category scores standard metrics more, hence associated with emoji use, while binary logistic regression was used to examine which category is statistically associated with emoji occurrence. Findings unveil that nouns dominate the corpus in normalized frequency (M = 0.675, SD = 0.161), followed by verbs (M = 0.083, SD = 0.100). However, verbs have the strongest influence of emoji use indicated by verb density (\b{eta} = 0.821, p = .001, 95% CI [0.332, 1.309]). The study concludes that in digital pragmatics of CA on X, emoji use association with lexical/pragmatic category can be explained by a hybrid approach of computational, statistical, and pragmatic methods, reflecting the interaction among machine learning, linguistic/lexical features, contextual representation, and pragmatic communication.
Mohammed Q. Shormani, Y. A. AlSohbani· 0 citations
This corpus-based study investigates the linguistic behavior and distributional patterns of three English near-synonymous verbs advise, suggest, and recommend. While these lexical items share a core denotational meaning, their actual usage is often governed by subtle nuances that pose challenges for EFL/ESL learners. Data were sampled from the Corpus of Contemporary American English (COCA), focusing on frequency distribution across eight distinct genres and an analysis of noun collocations using a Mutual Information (MI) score of 3 or higher to ensure statistical significance. The findings reveal that suggest is the most frequent and formal synonym, predominantly appearing in academic discourse to introduce research evidence and theoretical constructs. In contrast, advise and recommend occur more frequently in less formal digital genres, such as webpages and blogs, reflecting their practical and communicative functions. Collocational analysis shows that advise and recommend share a close semantic relationship, often co-occurring with professional advisers. However, advise is more strongly linked to personal interpersonal guidance and advice recipients, while recommend is heavily associated with official endorsements, medical treatments, and institutional bodies. These results demonstrate that the three verbs inhabit distinct distributional niches and are rarely interchangeable without affecting the register or tone of the discourse. This study underscores the importance of corpus-informed learning in helping non-native speakers navigate the subtle nuances and semantic preferences of English near-synonyms effectively.
Stress position in English words is well-known to be influenced by phonological, morphological, and (morpho)syntactic factors (esp. verbs vs. nouns). All these influences, however, are far from categorical, and very often stress varies from one word to another in a seemingly unpredictable manner. What is more, recent corpus-based work points to the relevance of opaque, i.e. etymological prefixes. It is, however, unclear why effects of opaque morphology should be productive synchronically. Moreover, there is an increasing body of research indicating that stress assignment works on the basis of lower-level information, i.e. analogy. In this paper we analyze the results of a pseudo-word experiment testing the production of trisyllabic verbs. Structural predictors tested in the analysis comprise syllable weight and prefixation (measuring the transparency of prefixes). To test analogy, we had a computational analogical model predict stress in the pseudo-words based on all verbs in the
Cambridge Pronouncing Dictionary
. Lexical support measures were then used as predictors alongside weight and prefixation in a mixed-effects logistic regression model. The results show that all three variables are significant predictors of main stress in the experimental data. Lexical support is strongly correlated with actual productions, but AML underpredicts some of participants’ preferences. These are, however, explained by syllable weight and transparency of the prefix in our regression model. We conclude that properties on different levels of representation are relevant for the prediction of stress position in English pseudo-verbs.