Many historical handwritten records in low-resource languages remain difficult to access through modern digital systems. This limits efforts to preserve and study cultural heritage at scale. Chagatai manuscripts exemplify these challenges within the Eastern Turki tradition. For centuries, it served as a major written language across Central Asia and supported a rich literary tradition. Large collections of Chagatai manuscripts still survive today, yet only a small amount of this material exists in digital form. As the technical literature specifically focused on Chagatai-HTR remains in its nascent stage, this review synthesizes indirect evidence from taxonomically related Perso-Arabic scripts to establish a foundational research framework. This article presents a systematic literature review following the PRISMA guidelines to examine artificial intelligence methods for handwritten text recognition (HTR) and text restoration in low-resource languages. Analyzing 50 studies published between 2020 and 2026, the review categorizes research trends into handwritten text recognition (HTR), optical character recognition (OCR), script classification, dataset development, and multimodal vision–language systems. The findings reveal a significant architectural shift from traditional segmentation-based CNN and RNN models toward transformer architectures and multimodal approaches. However, for Chagatai specifically, the primary obstacle is not the lack of advanced models but a critical scarcity of basic research infrastructure, including expert-verified transcriptions, annotation standards, and open benchmark datasets. Consequently, this article proposes a concrete development roadmap focusing on systematic digitization, expert annotation, transfer learning, and the creation of baseline models to enable reproducible evaluations.
Zhanibek Balabayev, S. Biloshchytska, Beibit Abdikenov et al.· Information· 0 citations
Neural machine translation (NMT) systems are widely used, but their performance remains strongly dependent on the availability of large-scale digital corpora, making translation for low-resource languages a persistent challenge. In parallel, large language models (LLMs) have recently emerged as a promising paradigm for multilingual text generation and translation; however, their behavior in low-resource settings remains largely underexplored. The challenge becomes even more acute for historical languages. Chagatai, a historical Turkic literary language of Central Asia with no native speakers, unstable orthography, and parallel data, represents an extreme case of such a condition. This study investigates whether transliteration significantly affects translation performance and how LLM-based and NMT-based systems compare under an extremely low-resource setting. To address these questions, we evaluated four source-text configurations (original Arabic script, expert manual transliteration, LLM-based transliteration, and rule-based Uroman transliteration) for translation into six target languages: Kazakh, English, Uzbek, Uyghur, Turkish, Russian, and Arabic. The results show that manual transliteration consistently yields the best translation performance, while noisy automatic romanization reduces these gains. For model comparison, GPT-4o was assessed alongside two fine-tuned NMT baselines, NLLB and TranslateGemma. The findings further show that LLM-based translation can be competitive with, and in some settings outperform, fine-tuned NMT systems, although this advantage comes with lower interpretability. Overall, these findings show that, for extremely low-resource historical languages written in non-Latin scripts, source-side representation is a decisive factor and may be as important as the choice of translation model itself.
A. Mansurova, Meruert Bekmukhamedova, Bekarys Baibolat et al.· Electronics· 0 citations