Sep 2026· Digital Technologies Research and Applications· Vol 5, pp. 122-131· 0 citations· 28 references
TL;DR
This study highlights the importance of integrating in-depth linguistic analysis, contextual semantic modelling, and cultural awareness into natural language processing-based translation systems by combining traditional linguistic insights with computational methods.
Abstract
Translation is defined as the process of transferring meaning from one language to another. It is an extremely difficult and complex process because it involves not only transferring words but also ideas, culture, linguistic customs, and meanings derived from syntactic elements, their arrangement, word structure, and derivation. This is especially true in languages with complex structures, such as Arabic, which is characterized by its multiple linguistic contexts. These contexts have not been adequately addressed by NLP (Natural Language Processing) applications in machine translation models due to the lack of diverse contexts where metaphor, figurative language, grammatical inflections, morphological patterns and their connotations, and sentence structure all play pivotal roles in determining meaning. Furthermore, spoken language, with its inherent phonetic and expressive characteristics, conveys the text into broader semantic spaces. These spaces are influenced by the effect of intonation on specific syllables, the speaker's psychological state, the listener's mood, and accompanying body language, which transforms meaning into other subtle details. All of this, and more, is absent from machine translation, no matter how hard its creators try to imbue it with human emotions and feelings. This study highlights the importance of integrating in-depth linguistic analysis, contextual semantic modelling, and cultural awareness into natural language processing-based translation systems. By combining traditional linguistic insights with computational methods, the research offers a framework that can contribute to improving the accuracy of machine translation from Arabic to English.
Machine Translation has become one of the major application areas of Artificial Intelligence (AI) and Natural Language Processing (NLP), especially in multilingual countries like India. Although recent Neural Machine Translation systems have shown good performance for several language pairs, translation quality is still inconsistent for many Indian languages because of linguistic and structural differences between English and Indian language families. Most Indian languages are morphologically rich and contain flexible word order, complex agreement patterns, compound constructions, and context-dependent grammatical forms. Because of this, direct translation from English often produces structurally incorrect or semantically weak output. In many existing systems, the source sentence is passed to the translation model without sufficient linguistic analysis. As a result, ambiguity present in the source text propagates further during translation. This work focuses on the importance of linguistic enrichment before the translation stage. The proposed framework, named Unified Linguistic-Aware Pre-Parsing Framework, introduces a coordinated pre-processing layer for English-to-Indian Machine Translation (MT). A key contribution of this research is the development of a novel linguistically enriched intermediate representation that extends beyond conventional text normalization. By transforming noisy input text into linguistically enriched translation-ready representation, the proposed approach facilitates effective knowledge transfer to machine translation models, leading to improve contextual adequacy, linguistic fidelity, and overall translation performance. The framework combines multiple linguistic processing stages including POS tagging, NE detection, clause boundary analysis, contextual token handling, syntactic structure preparation, and morphology-related processing. Instead of executing these modules independently, the proposed system allows interaction between lexical, syntactic, and morphological information during analysis. This helps reduce structural ambiguity and improves sentence-level interpretation before translation begins. The need for such a framework becomes more relevant in the context of Indian languages where morphology and grammatical relations carry significant semantic information. This framework is especially relevant for Indian languages, where semantic information is often encoded through morphological variations and grammatical dependencies. The proposed framework can be effectively integrated with both conventional machine translation architectures and modern large language models. The overall study highlights how classical linguistic analysis can still play an important role in improving multilingual AI systems for Indian languages.
Prashant Chaudhary, Pavan Kurariya, Jahnavi Bodhankar et al.· NLP & Big Data· 0 citations
Text which has been translated from another language tends to carry with it evidence of translation$\unicode{x2014}$hence, it is often referred to as $\textit{translationese}$. Multilingual large language models (MLLMs) generate text in a variety of languages. However, it is still unclear if MLLMs'generations resemble internal translation (from English or, potentially, other languages) and, thus, result in translationese. Here, we ask the following research questions: (1) Does text generated by MLLMs resemble translationese? (2) How does translationese produced by MLLMs differ from translationese produced through direct translation? We leverage established indicators of translated text to evaluate text generated by state-of-the-art MLLMs in five languages, comparing to both non-translated and human-written baselines in order to isolate translationese from other kinds of interference. Through the use of high-accuracy classification models, analyses of variance on individual linguistic features, and the collection of human annotations in a subset of two languages (German and Spanish), we assess the translationese content of MLLM generations and examine the key features that distinguish MLLM-generated text from typical translation-related interference.
Maria R. Valentini, Téa Wright, Julisa Granados et al.· 0 citations
This paper analyses various recent state-of-the-art variants of large language models (LLMs) and neural machine translation (NMT) for Indian languages in comparison to statistical machine translation (SMT) and tackles key questions, such as idiomatic expressions, morphologically complex grammar or the scarceness of parallel corpora.
Jayanand A. Kamble, S. Jadhav, V. J. Kadam· International Journal of Inf...· 0 citations
Ambiguity is widely recognized as one of the fundamental challenges in natural language processing. It significantly affects both human language comprehension and computational language interpretation. Among its various forms, structural and textual ambiguity present particular difficulties for linguistic analysis and artificial intelligence systems. This study investigates the linguistic characteristics of structural and textual ambiguity in English and examines how contemporary language models process and interpret ambiguous linguistic structures.The research is based on a corpus of selected a three authentic English sentences containing structural and textual ambiguity. An experimental study was conducted with 45 students at Azerbaijan University of Languages. In the first phase, structurally ambiguous sentences were analyzed using the PRAAT computer program to examine their linguistic and prosodic features. In the second phase, students were instructed to generate texts in ChatGPT using the selected structurally ambiguous sentences as prompts. The resulting texts were subsequently analyzed and compared with outputs generated by ChatGPT, Gemini and Claude to investigate how these artificial intelligence systems interpret, contextualize and resolve structural ambiguity at the textual level.The findings provide insights into the mechanisms employed by large language models in processing ambiguous language and demonstrate similarities and differences in their contextual interpretation of structurally ambiguous constructions. The study contributes to the theoretical understanding of structural and textual ambiguity highlighting their communicative functions in linguistics, natural language processing and artificial intelligence.The relevance of the study lies in its interdisciplinary approach, integrating theoretical linguistics, experimental research, and AI-based language analysis. The results contribute to the growing body of research on ambiguity resolution and offer practical implications for the development of more accurate and context-sensitive natural language processing systems.
Khanbutayeva Leyla Musa· Aposta: Revista de Ciencias...· 0 citations
The subject of the research is the semantic, grammatical, and pragmatic characteristics of translations generated by DeepSeek in comparison with DeepL and Google Translate. The object is machine translation in the Russian–Chinese language pair using generative language models. The relevance is determined by the contradiction between the expanding use of large language models in translation and the insufficiently defined boundaries of their effectiveness with typologically and culturally distant languages. The aim is to determine the boundaries of DeepSeek's effective application in both translation directions. The objectives include characterizing the model's technological features, comparing cognitive mechanisms of language processing by humans and artificial intelligence, and empirically testing translation quality against DeepL and Google Translate according to semantic accuracy, grammatical correctness, and pragmatic adequacy. The study employed comparative analysis, cognitive modeling, and interpretive analysis on a corpus of 30 phraseological and culturally marked units. The scientific novelty lies in the systematization of knowledge about generative neural networks in translation theory and in the comparative assessment of three systems on a unified corpus according to three complementary criteria. The author's contribution consists in refining the understanding of similarities and differences between human and machine cognitive mechanisms. The main findings are as follows: DeepSeek outperforms DeepL and Google Translate in conveying idioms and cultural realia, however its functional adaptation may lead to semantic shifts, necessitating professional post-editing in terminologically dense texts. The most justified application is producing draft translations and finding contextual equivalents, while final verification should remain with the human translator.
Ilia Alekseevich Konstantinov· Филология научные исследован...· 0 citations