Aug 2026· Theory and Practice in Language Studies· Vol 16, pp. 2670-2680· 0 citations
Abstract
Universal and multilingual large language models cannot correctly simplify Kazakh texts in accordance with internal linguistic patterns. Using LLM without a linguistically sound simplification model leads to uncontrolled generation and increases the risk of lexical and grammatical errors. Therefore, there is a need to create formalized linguistic patterns for Kazakh text simplification, which will facilitate the creation of automatic simplification systems within LLM. This can be achieved through the analysis of manually adapted materials. A corpus of 20 Kazakh texts was annotated by the authors for the analysis, each presented in three versions such as the original (O) and two successive adaptation levels (A1 and A2). At the A1 stage, syntactic transformations predominate, while A2 demonstrates lexical stabilization. The text size is comparable at all levels, which ensures control and reliability of the analysis. The process is formalized through a two-level tagging system: lexical (LS) and syntactic simplification (SS) that records each operation. Frequency analysis revealed a limited core of dominant techniques, supplemented by peripheral layers, forming structured polyoperational architecture. Transitions from A1 to A2 demonstrate a process shift toward lexical changes while maintaining discursive coherence. The results emphasize the importance of linguistic expertise, operation typologies, and manual annotation for the interpretability and quality of systems. Hybrid approaches with humans in the loop improve the accuracy and reliability of ATS.
It was concluded that texts generated by artificial intelligence constitute a separate linguistic phenomenon with its own set of characteristics, which requires a special typology and a flexible, updatable analysis methodology.
L. Kravets, Viktória Stefuca, N. Libak et al.· Computer and Decision Making...· 0 citations
This paper is focused on the problem of overcoming macro-typological asymmetry in translating texts from Russian into Arabic and from Arabic into Russian. The paper is dedicated to a comprehensive study of the mechanisms of cross-linguistic meaning transfer in cases where direct transfer (transcoding) proves impossible due to structural interference. Particular attention is paid to the deconstruction of complex verbal phrases, metaphorical expressions, and culture-specific vocabulary in the absence of direct dictionary equivalents. The aim of the research is to develop a comprehensive formal model of linguistic-semantic and syntactic transformations for the Russian-Arabic language pair based on dependency grammars, utilizing the empirical data of parallel text corpora. The scientific novelty of the study lies in the fact that, for the first time, the mechanisms for overcoming lexical-syntactic lacunarity between typologically distant languages (such as demetaphorisation, cross-part-of-speech transposition (recategorization), and syntactic reconfiguration) are presented in the form of computable syntactic dependency trees based on a parallel macrotext. Furthermore, the applied potential of Russian-Arabic parallel arrays is substantiated for the first time by applying the principle of corpus triangulation to identify translation patterns. As a result, the key text adaptation strategies (nominal, semantic, and explicative) were identified and systematized. The results obtained showed that achieving pragmatic equivalence is algorithmically impossible without the structural expansion or reduction of the source graph, and that the use of corpora safeguards the translator from subjective deviations and the phenomenon of "translationese".
Dawood Jaafar Al, Borisovna Kozerenko Elena· Philology. Theory & Prac...· 0 citations
This article examines the theoretical and methodological foundations for identifying the grammatical minimum to be included in a lexico-grammatical dictionary designed for teaching Kazakh as a foreign language. The study is grounded in the communicative needs of foreign language learners (non-native speakers) and adopts frequency, structural simplicity, and cultural-cognitive significance as the key criteria for selecting grammatical material. The grammatical categories of the main parts of speech in Kazakh (nouns, verbs, adjectives, numerals, pronouns, adverbs, particles, and interjections) are systematized and presented through a contrastive analysis with English. Grammatical forms are introduced not through direct translation but via functional explanation and descriptive equivalents. Grammatical units that lack direct equivalents are explained through examples that reveal their semantic and pragmatic functions. The results demonstrate that presenting lexis and grammar as an integrated system enhances learners’ communicative competence and increases the practical value of the dictionary. The article provides a theoretical foundation for instructional and lexicographic works aimed at teaching agglutinative languages as foreign languages.
A. Soltanbekova, T. Ramazanov, Q. B. Slyambekov et al.· Theory and Practice in Langu...· 0 citations
This study highlights the importance of integrating in-depth linguistic analysis, contextual semantic modelling, and cultural awareness into natural language processing-based translation systems by combining traditional linguistic insights with computational methods.
Hilal Abdul-Raziq Sadiq, Zaxid Maxmudovich Islamov, R. Matibaeva et al.· Digital Technologies Researc...· 0 citations
In this paper, we are exclusively concerned with the part of grammar that deals with the structure of sentences. This is called syntax. Not only the grammatical units of language were explored but the division of selected sentences into constituents (units) was also analyzed. To achieve this feat, sentences were separated into words and finally, words were regrouped on the basis of relationship between them. This paper has gone further to explain how the (agent) or subject of a sentence is identified through grammatical units. The grammatical units were introduced on the hierarchical order [down-up]. The general syntactic framework we have adopted is inspired by the theories of language developed by Noam Chomsky. The choice of this study is based on the assumption that English and French are “closely related and well documented languages” (Tanja et all, 2010:110) and the two (duo) “constitute a minimal pair suitable for micro-comparison (Kayne, 2005). Most learners that were not always comfortable with the syntax and structure of the two languages would be familiar with their components and syntax of English and French languages armed with this work.
T. A. Balogun, G.S. Idowu· Eureka-Unilag· 0 citations
Ambiguity is widely recognized as one of the fundamental challenges in natural language processing. It significantly affects both human language comprehension and computational language interpretation. Among its various forms, structural and textual ambiguity present particular difficulties for linguistic analysis and artificial intelligence systems. This study investigates the linguistic characteristics of structural and textual ambiguity in English and examines how contemporary language models process and interpret ambiguous linguistic structures.The research is based on a corpus of selected a three authentic English sentences containing structural and textual ambiguity. An experimental study was conducted with 45 students at Azerbaijan University of Languages. In the first phase, structurally ambiguous sentences were analyzed using the PRAAT computer program to examine their linguistic and prosodic features. In the second phase, students were instructed to generate texts in ChatGPT using the selected structurally ambiguous sentences as prompts. The resulting texts were subsequently analyzed and compared with outputs generated by ChatGPT, Gemini and Claude to investigate how these artificial intelligence systems interpret, contextualize and resolve structural ambiguity at the textual level.The findings provide insights into the mechanisms employed by large language models in processing ambiguous language and demonstrate similarities and differences in their contextual interpretation of structurally ambiguous constructions. The study contributes to the theoretical understanding of structural and textual ambiguity highlighting their communicative functions in linguistics, natural language processing and artificial intelligence.The relevance of the study lies in its interdisciplinary approach, integrating theoretical linguistics, experimental research, and AI-based language analysis. The results contribute to the growing body of research on ambiguity resolution and offer practical implications for the development of more accurate and context-sensitive natural language processing systems.
Khanbutayeva Leyla Musa· Aposta: Revista de Ciencias...· 0 citations