Skip to content
Review Open access

Advancing Large Language Models for Low-Resource Languages: A Systematic Review of Pretraining, Adaptation, and Ethical Challenges

2026 · Computer Modeling in Engineering & Sciences · 0 citations · 144 references

TL;DR

This systematic review examines recent progress in the pretraining and adaptation of LLMs for Low-Resource Languages (LRLs) and focuses on the ethics in AI practice, the development of corpora through communities, and interdisciplinary research collaboration among computational linguists, social scientists, and digital humanists.

Abstract

: In recent years, the rapid advancement of Large Language Models (LLMs) has significantly transformed natural language processing (NLP), enabling impressive performance across a wide range of tasks. However, these developments have largely benefited high-resource languages, leaving many low-resource and underrepresented languages at risk of further digital marginalization. Addressing this imbalance is crucial to building more inclusive and culturally sustainable AI systems, which is motivating growing research interest in adapting LLMs for linguistically diverse and resource-scarce communities. This systematic review examines recent progress (2020–2025) in the pretraining and adaptation of LLMs for Low-Resource Languages (LRLs). Analysed 812 records obtained in the large databases and using PRISMA criteria, 140 core studies were identified. The innovations in data augmentation and parameter-efficient fine-tuning approaches can be outlined in this selection process. It combines major innovations on data-driven augmentation, parameter-efficient fine-tuning and morphologically rich and underrepresented language script-sensitive tokenization. The results highlight the growing effectiveness of culturally aware standards such as IrokoBench and BLEnD and show that approaches to lightweight adaptation eliminate high computational costs while maintaining language accuracy. The review focuses on the ethics in AI practice, the development of corpora through communities, and interdisciplinary research collaboration among computational linguists, social scientists, and digital humanists. The task of generating a diversified dataset, typology-conscious modelling strategies, and open-source multilingual benchmarks should be prioritized in future research as one possible solution to the existing digital language gap worldwide.

Read PDF

Similar papers

Review Open access Aug 2026

Bridging the AI Language Divide: A Systematic Review of NMT and LLMs in Low-Resource Translation

The rapid evolution of AI-driven language technologies has inadvertently widened the gap between high-resource and marginalised languages. Despite significant progress in AI-driven translation for high-resource languages, low-resource languages remain underrepresented due to limited data, a lack of benchmarks, and evaluation challenges. This study presents a comprehensive systematic review of machine translation for low-resource languages, focusing on advances in neural machine translation (NMT) and large language models (LLMs) between 2017 and 2025. Following Preferred Reporting Items for Systematic Reviews and Meta-Analyses (PRISMA) guidelines, 63 studies were selected from the 1696 articles in the Scopus, Web of Science, and Google Scholar databases. The review identifies five dominant methodological approaches: data augmentation, back-translation, transfer learning, pre-training, and parameter-efficient fine-tuning. The findings reveal that model performance is highly dependent on resource availability: transformer-based NMT excels in moderate data settings, while LLMs demonstrate promising zero-shot and few-shot capabilities in extremely low-resource scenarios. Hybrid NMT–LLM approaches emerge as a particularly effective paradigm. The study also highlights critical challenges, including the absence of standardised benchmarks, over-reliance on inadequate evaluation metrics such as Bilingual Evaluation Understudy (BLEU), limited human evaluation, and significant geographic and linguistic underrepresentation. Additionally, ethical concerns related to bias, cultural representation, and community engagement are increasingly relevant. The findings contribute to advancing inclusive and equitable AI-driven language technologies.

Sweet Agrawal, A. Agbeyangi · 0 citations
Open access 2026

Challenging the Abilities of Large Language Models in Italian: a Community Initiative

CALAMITA is conceived as a rolling benchmark, enabling continuous integration of new tasks and models, and argues that this combination offers a blueprint for other languages and communities seeking inclusive and rigorous LLM evaluation practices.

Malvina Nissim, Danilo Croce, V. Patti et al. · 0 citations
Preprint Jul 2026

PluraMath: Extending Mathematical Reasoning Evaluation Beyond High-Resource Languages

Mathematical reasoning has become a central task for evaluating and tuning reasoning Large Language Models (LLMs), yet existing benchmarks remain heavily biased toward high-resource languages, with English and Chinese dominating both pre-training corpora and evaluation suites. The recently released PolyMath (Wang et al., 2025) dataset represents a significant step forward, yet its coverage is still limited to 18 only high-resource languages. To address this gap, we introduce PluraMath, an extension of PolyMath to 18 additional {underrepresented languages spanning 6 language families -- ranging from mid-resource to extreme low-resource settings. We constructed the dataset through a human-curated pipeline, where native speakers thoroughly validated pre-computed translations. Using PluraMath, we then benchmark 27 reasoning LLMs across four model scales -- small, mid-size, large, and closed-source ensembles -- probing the multilingual mathematical reasoning capabilities of state-of-the-art models under diverse linguistic conditions. Our fine-grained analysis confirms a persistent gap in mathematical reasoning performance between high-resource and underrepresented languages, with stronger results largely associated with better instruction-following ability. We fully open-source our dataset, data acquisition pipeline, and evaluation framework, with the goal of lowering the barrier to multilingual benchmark development for underrepresented communities.

Daryna Dementieva, N. Babakov, Kathy Hammerl et al. · 0 citations
Open access Aug 2026

Harnessing Advanced Transfer Learning Techniques in GPT-2 for Real-World Multilingual Applications

: In an era of increasing demand for robust multilingual natural language processing, leveraging advanced transfer learning techniques has become essential. This paper explores the application of the GPT-2 model using a comprehensive Serbian dataset of 750 million tokens. By employing meticulous data preprocessing, effective tokenization, and precise hyperparameter optimization with Optuna, the model's performance in language tasks is significantly improved. These findings underscore the model's adaptability to diverse linguistic contexts, facilitating deployment in real-world applications. The significant performance improvements highlight broader applicability in multilingual AI environments. The paper addresses key challenges such as data heterogeneity and computational efficiency, providing insights and proposing strategies for future research. By overcoming these challenges, the research demonstrates the transformative potential of refined GPT-2 models in multilingual AI. The advancements made lay a solid foundation for further exploration and refinement of multilingual language models, paving the way for more inclusive and accurate AI-driven communication tools.

Dejan Dodi, Ć. DušanREGODI, Ć. AnaVUKI et al. · 0 citations

Have Large Language Models Improved Research Methodology?

Whether contemporary LLMs can reproduce the research outcomes of a fully documented human study: a 1991 article that identified dermatophytosis (ringworm) in historical fine art was evaluated.

Fredric Narcross, Robert Marks · 0 citations
Review Open access Jul 2026

A Systematic Review of Transfer Learning and Data Augmentation in Neural Machine Translation of Low-Resource Languages

The findings show that a hybrid combination of TL and DA with a self-supervised objective is the most effective solution for extremely low-resource scenarios, capable of producing the highest translation quality and outperforming baseline models and traditional methods such as Statistical Machine Translation (SMT).

Nur Fikri Khuluq, Muhammad Naufal Muzhaffar, Shofwatul Uyun · 0 citations