Aug 2026· Engineering, Technology & Applied Science Research· 0 citations· 13 references
TL;DR
The reproducible engineering process for building a ground truth-validated dataset for deteriorated Old Sundanese manuscripts from the Geusan Ulun Museum in Indonesia is described and can be applied to low-resource historical scripts, connecting cultural heritage and computational modeling.
Abstract
While cultural preservation and the digitization of historical manuscripts have a long history, low-resource writing systems continue to be poorly represented due to the absence of reliable ground truth datasets. This paper describes a reproducible engineering process for building a ground truth-validated dataset for deteriorated Old Sundanese manuscripts from the Geusan Ulun Museum in Indonesia. The Old Sundanese Manuscripts Dataset (OSMD) was constructed by combining controlled digitization, contrast enhancement, edge-based object extraction, morphological image processing, expert annotation, and optimized dataset packaging. The dataset contains 1,300 word-level samples for five consonant classes: Ha, Na, Ca, Ra, and Ka. To assess the potential of the dataset, baseline classification experiments were conducted using HOG+SVM and a lightweight CNN, where the latter outperformed (p<0.01) with 0.842 accuracy and 0.833 macro F1-score. The absence of contrast enhancement resulted in a notable drop in performance, and cross-validation demonstrated low variance (±0.006). The process followed in this study can be applied to low-resource historical scripts, connecting cultural heritage and computational modeling.
Palimpsests are ancient manuscripts in which the original text has been partially erased and overwritten by a later one, which makes them very difficult to decipher. This study examines the possibilities for analyzing, evaluating, and classifying the legibility of such manuscripts using modern image processing and mach...
G. Dimitrov, P. Tsvetkova, Pavel S. Petrov et al.· Automation, Control, and Inf...· 0 citations
This paper is devoted to the development of a unified alphabet of Ancient Turkic dialects for solving the problems of automated recognition of Orkhon–Yenisei runic inscriptions. The relevance of the study is determined by the fragmentation of existing rune classifications and the absence of a unified system for correlatin...
A. D. Borodina, M. Kyzyl-ool, R. Kochkarov· Digital Solutions and Artifi...· 0 citations
Pottery is a primary source for reconstructing the chronological and economic dimensions of past societies. Archaeologists often document ceramic finds through technical drawings and handwritten metadata. This metadata is critical for dating, provenance attribution, and cross-site comparison, but remains inaccessible t...
Gissu Valentina Naghavi, Dominik Hagmann, M. Kampel et al.· 0 citations
Purpose: The digitization of historical documents presents fundamental challenges for modern information retrieval and Artificial Intelligence (AI) systems. Optical character recognition (OCR) errors in source corpora propagate through retrieval-augmented generation (RAG) pipelines, compromising the factual accuracy of...
Marina Gómez Rey, Patricia Callejo, M. Muñoz-Organero et al.· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.