Aug 2026· Automatic Documentation and Mathematical Linguistics· Vol 60, pp. S225 - S230· 0 citations· 18 references
TL;DR
The experimental results revealed the strengths and weaknesses of various approaches to subword segmentation and identified the most effective tokenization strategies under the conditions of the morphological complexity of the Tajik language.
The article reports on a corpus-driven analysis of how Sinclairian extended units of meaning (EUMs) are treated in English into Slovene translation. The main goal of the underlying research was to test to what extent the EUM is perceived as the proper (extended) unit of translation as proposed by Zethsen (2021). The point of departure was a list of high-frequency EUMs that have previously been identified in a monolingual English corpus. Subsequently, these EUMs were checked in the OpenSubtitles 2018 parallel English-Slovene corpus. Although the overall number of EUMs was lower than expected, the results indicate that the rate of successful translation of EUMs in real-life circumstances is in the 30–40% range, while in the remaining majority of translations semantic prosody is either omitted or it undergoes explicitation by means of denotation. This low rate should encourage authors in the field of corpus-based translation studies to spread and raise awareness of semantic prosody. Another implication of the study in the field of higher education of translators is the need to alert translation students to this elusive phenomenon.
P. Jurko· Across Languages and Culture...· 0 citations
The article is devoted to a comprehensive lexical analysis of M. Simonyan’s English-Armenian dictionary, “Small Dictionary from the English Language, Translated into the Armenian Dialect” (Madras, 1803). Using the example of the first English-Armenian dictionary, the study analyzes the qualitative and structural features of the target language—Armenian. Following a brief presentation of the historical context, the primary focus shifts to the manifestations of linguistic variability, orthographic irregularities, and thematic groups present in the dictionary. The rich dialectal layer of the dictionary and the use of vernacular synonyms undergo special examination. Through a comparative analysis with the Old and New Haigazian dictionaries, M. Simonyan’s terminological and neological efforts (morphological, derivational, and semantic coinages, particularly in the realm of rhetorical terms) are highlighted. The unique value of this work during the transitional phase of Armenian lexicography is demonstrated.
Վանուհի Բաղրամյան· Language and Linguistics· 0 citations
Ayvazyan’s textbook-manual “The Orthography of Our Armenian Language”, in addition to its primary pedagogical function, also reflects the author’s grammatical understanding and his approaches to various linguistic issues.
The work examines numerous important questions related to the invention of the Armenian alphabet, problems of letter pronunciation, as well as matters of transliteration and punctuation. Ayvazyan provides comprehensive knowledge about the prevailing views on the creation of the alphabet and their relationship to the letters of other languages. In accordance with the demands of the time, the author devotes significant attention to issues concerning transliteration, setting forth all the principles that translators had to follow in order to make their work as uniform as possible.
Ayvazyan’s grammatical approaches are based on the rich experience of the historical development of the Armenian language. The author was able to accurately foresee numerous orthographic and punctuation functions used in Armenian language (Ashkharhabar), most of which remain relevant in modern Armenian. Moreover, the parallels and comparisons drawn with European languages made the author’s work more comprehensive and an exceptionally effective grammatical study for its time.
Hasmik Iritsyan· Journal of Armenian studies· 0 citations
The article describes and analyzes the comparative meaning of the polysemantic lexeme
olchaan
‘exactly the same as’, ‘just like’ in the Tuvan language. The research methodology is based on the semantic approach to the comparative description of comparative constructions in the Ural-Altaic languages of Siberia, developed at the Institute of Philology of the SB RAS. This approach distinguishes between the content plane (comparata, parameter base, parameter aspect, exponent, relation) and the expression plane. The material includes selection of examples from fiction and folklore texts, as well as the author’s subjective observations of Tuvan colloquial speech. It has been established that
olchaan
expresses similative relations, indicating the maximum degree of similarity. This lexeme can function as an adverbial modifier, an attribute, and, most characteristically, as a nominal predicate. As a present-tense predicate,
olchaan
does not take person-number affixes but combines with the evidential affirmative particle
-tyr
and the auxiliary verb
bol-
‘to be, to become’ to express tense forms. Two main types of constructions are identified: with an explicit and implicit basis for the comparative parameter. Constructions with this word are characterized by a high degree of contextual dependence, regularly realized in the ellipsis of the object of comparison. The lexeme
kara
‘black’, as a semantic intensifier, gives the comparison an expressive nuance. The word
olchaan
occupies a special place in the system of comparative means in the Tuvan language, combining the functions of a comparison marker and a predicate, and participates in the formation of the figurativeness of an utterance.
Contemporary linguistics shows a growing interest in studying language as a polysystem, with particular attention to the interrelation of language, thought, and the cultural worldview. Within this framework, the analysis of literary concepts that represent collective traumatic experiences becomes particularly relevant. This research aims to identify and describe the distinctive linguostylistic features used to represent the concept of “hunger” (golod) in the Russian-language segment of Kazakh prose of the 1930s. The study is based on an integrative approach combining methods of corpus linguistics (frequency, concordance, and collocational analysis employing statistical metrics such as t-score and Mutual Information), cognitive-discursive analysis for reconstructing mental models. A specialised text corpus of works by the key authors of the period (S. Mukanov, B. Mailin, G. Musrepov) was compiled. The findings reveal that the concept of “hunger” (golod) is represented through a stable set of linguostylistic devices. The lexical core of the semantic field is formed by the nominatives “bread” (khleb) (13.6 per 10,000 words), “hunger” (golod) (12.3), “death” (smert’) (9.3), and “to eat” (est’) (10.4), which actualise the biological, social, and symbolic dimensions of the tragedy. Collocational analysis confirmed stable associations with adjectives such as “frosty” (moroznyĭ) (MI = 8.7). Distinct authorial strategies were identified: an emphasis on social evil (Mukanov), natural force (Mailin), and redemption (Musrepov). The results confirm the hypothesis that linguostylistic devices constitute the principal mechanism for constructing literary meaning and emotional impact, reflecting the collective experience of dehumanisation and the loss of traditional ways of being.
M. Jakypbekova, Shara Kyyakhmetova, Zhanar Seisembayeva et al.· Theory and Practice in Langu...· 0 citations
Talking openly about anxiety and depression (A&D) remains difficult for many people because of the stigma surrounding mental illness. Anonymous online platforms such as Reddit provide a space where users can express their thoughts and emotions more freely. This study considers how individuals linguistically construct and intensify emotional distress by examining (1) the adjectives used to express A&D, (2) the content-word collocates that co-occur with these adjectives, (3) the lexical features of these collocates, and (4) the emotional meanings conveyed through these collocational patterns. The dataset consisted of 1,440 Reddit posts (approximately 300,000 tokens) systematically sampled from the r/Anxiety and r/Depression subreddits between 2023 and 2025. An observational mixed-methods corpus linguistic approach was used to examine the data. Quantitative corpus linguistic analyses were carried out using AntConc, including frequency profiling and Mutual Information (MI) analysis, and were enhanced by qualitative concordance and Keyword-in-Context (KWIC) analysis to examine collocational patterns in context. The analysis shows a predominance of negatively valenced adjectives (e.g., anxious, depressed, hopeless, and suicidal), whose meanings are systematically intensified through their collocational environments. The collocates show distinctive lexical features. These include clinical nouns, linking and change-of-state verbs, and degree and frequency adverbs. These features construct varying levels of affective intensity and psychological distress. Emotional meaning is encoded in recurrent collocational patterns. Individual lexical items reveal only part of this meaning. This shows the value of collocational analysis for digital mental health research. The findings also possess practical implications. They may help improve the diagnostic sensitivity of automated digital mental health tools and foster more empathetic clinical communication.
Y. Yuhan, S. Lau, Handoko Handoko et al.· Jurnal Arbitrer· 0 citations