Skip to content
Open access

Good Enough to Read?

Aug 2026 · Journal of Arabic and Islamic studies · 0 citations

Abstract

This article proposes a pragmatic workflow for accelerating access to and analysis of large Arabic manuscript collections. Faced with the immense volume of historical texts and the inherent challenges of traditional approaches to hand-written text recognition (HTR), the proposed method combines off-the-shelf, cloud-based Optical Character Recognition (OCR) with large language models (LLMs). Despite the imperfect output often produced by Arabic HTR, this research demonstrates that LLMs can effectively derive meaningful summaries and key insights, enabling rapid triage for scholars. The workflow significantly enhances research efficiency by allowing for the swift identification of valuable documents, thereby freeing human expertise for in-depth analysis. While this approach holds particular importance for endangered archives, such as those in the Sudan, the paper points to its broader potential: manuscript scholars mastering AI tools not to replace but to support and enhance human understanding. Keywords: Arabic manuscripts • Handwritten text recognition (HTR) • Artificial intelligence • Digital humanities • Endangered archives, Sudan

Read PDF