A Comprehensive Analysis of Arabic Natural Language Processing Research: Trends, Topic Evolution, and Research Gaps -- A Bibliometric and Topic-Based Study
A bibliometric and topic-based analysis of 7,120 Arabic NLP papers published between 1960 and 2026, sourced from five platforms plus an additional targeted OpenAlex subset, and offers recommendations to prioritize under-resourced dialects and develop culturally aligned benchmarks.
Abstract
Arabic Natural Language Processing (NLP) has grown rapidly over the past decade, driven by digital transformation in the Arab world, social media, and large language models (LLMs). Despite this growth, a comprehensive quantitative meta-analysis remains absent. This study presents a bibliometric and topic-based analysis of 7,120 Arabic NLP papers published between 1960 and 2026, sourced from five platforms (arXiv, ACL Anthology, Semantic Scholar, Crossref, OpenAlex) plus an additional targeted OpenAlex subset. We employ BERTopic for topic modeling, regression analysis, social network analysis, and geographic mapping. Our findings show a significant publication surge after 2020, driven by transformer models and LLMs. Topic modeling identifies 19 themes, the largest centered on text, speech, translation, and recognition (2,942 papers). Citation analysis reveals a positive correlation between paper age and citations (r = 0.245, p<0.001); regression (R^2 = 0.105) shows that indexing in OpenAlex or Semantic Scholar and institutional affiliation are associated with higher citations. Saudi Arabia, the United States, and Egypt lead in research output. A task-dialect gap matrix identifies understudied areas, including summarization for Maghrebi, Iraqi, and Sudanese dialects. The largest topic has the highest H-index (90), followed by sentiment analysis (57). Our quantitative approach complements existing qualitative surveys and offers recommendations to prioritize under-resourced dialects and develop culturally aligned benchmarks.
While sentiment analysis has matured from an experimental technique to a core methodology in communication, it risks methodological stagnation due to data source limitations, and future research is suggested to focus on multimodal analysis and diverse digital platforms to overcome these constraints.
Sadettin Demirel· İletişim Kuram ve Araştırma...· 0 citations
An integrated framework based on transformer architecture for topic modeling and sentiment analysis for Hindi and Italian social-media discourse, customer reviews and news corpus is introduced and it is suggested that there is clear benefit for morphologically complex text and mixed script text for using contextual embeddings and language-specific pretraining.
Sunita Basalingayya, T. J. Peter· Journal of Intelligent Decis...· 0 citations
Machine translation is a language processing technology and cross-lingual intelligent task, which adopts computer algorithms and artificial intelligence to automatically transform text between different natural languages without manual sentence-by-sentence participation. In recent years, machine translation has been increasingly widely applied in translation practice, and relevant research has gradually become a popular academic hotspot. As a typical representative intelligent technological application in the fields of artificial intelligence and language services in China, machine translation is of vital importance to grasping the developmental context and evolutionary pattern of China’s artificial intelligence industry and language service sector, From the perspective of development trends, machine translation research is expected to remain a cutting-edge topic in translation studies. Based on 497 research papers published in CNKI during the research period, this paper adopts the CiteSpace information visualization software. From five dimensions including authors, research institutions, burst keywords, keyword relevance and timeline, it focuses on research hotspots and development trends, conducts a systematic analysis of domestic literature in the field of machine translation in China, and summarizes the research focuses and frontier trends in this domain. Using CiteSpace as the visualization analysis tool, this study finds that domestic machine translation research from 2020 to 2025 presents the following characteristics: it is led by core authors yet lacks adequate team collaboration; foreign language universities and top research institutions occupy a dominant position, and regional features have formed distinctive research clusters. The main research line remains stable, the research on human-machine relationship keeps advancing, and interdisciplinary features are becoming increasingly prominent. Research hotspots have been upgraded in technology, application scenarios and translation quality, forming a research pattern featuring the linkage of technology, process and data as well as the in-depth integration of industry, academia and research. Meanwhile, this study also predicts the potential development trends of machine translation in such directions as intelligent human-machine collaborative translation, multimodal translation, and the application of large language model-based translation.
Jun-Yao Yu, Hongyan Liao· International Journal of Eng...· 0 citations
Poetry is a unique form of expression valued for its role in preserving cultural heritage. Analyzing Arabic poetry is time-consuming and requires a high level of linguistic expertise; therefore, computational methods are useful, as they enable large-scale, extensive, and systematic analysis of poetry, thereby improving its accessibility for researchers and students. This article presents the first systematic review of natural language processing (NLP) and machine learning (ML) approaches for Arabic poetry. It addresses the question of which research tasks, methodologies, datasets, and evaluation approaches have been applied to Arabic poetry, and which trends and research gaps can be identified in the existing literature. In accordance with Preferred Reporting Items for Systematic Reviews and Meta-Analyses (PRISMA), the author conducted an exhaustive search across six major academic databases (ACL Anthology, IEEE Xplore, ACM, SpringerLink, Science Direct, and Google Scholar) for relevant studies published between January 2010 and May 2025. Eligibility was evaluated in several phases, and re-examination was conducted to ensure accuracy. The author performed task-level categorization, extracted key characteristics from each study, synthesized the findings, and presented them in tables and figures to highlight the main trends and research gaps in the literature. This study presents the first structured task-level synthesis of the field, identifying methodological trends, detecting evaluation inconsistencies, and highlighting research gaps that have not been critically consolidated before. Furthermore, the author assembled a comprehensive collection of available datasets and resources to promote standardized assessment.
The rapid proliferation of hotel reviews on online travel platforms, such as Booking.com and TripAdvisor, has necessitated the processing of large volumes of textual data for tourism researchers. The majority of existing Sentiment Analysis (SA) studies are limited to single-model approaches, which often lack validation against human-annotated labels and are restricted to document-level analysis. In this study, a structured multi-model framework is proposed for Aspect-Based Sentiment Analysis (ABSA) applied to 56,790 English hotel reviews collected from 30 hotels in Northern and Southern Cyprus. Three distinct generations of AI models are systematically compared: VADER (rule-based), ChatGPT/GPT-4o-mini (Large Language Model, prompt engineering), and DistilBERT (transformer). Eight service quality themes—location, food, service, cleanliness, rooms, facilities, value, and atmosphere—are extracted by generating structured JSON outputs via GPT-4o-mini, and the results are validated against human-annotated labels (achieving 91% agreement on a sub-sample of 100 reviews). The proposed framework demonstrates a scalable and reproducible methodology for aspect-oriented sentiment analysis in the hospitality domain.
Ayçin Giritli, Salahi Erensel, Ali Ozturen· Annual International Compute...· 0 citations