Skip to content
Open access

Towards the Automated Recognition of Ancient Sundanese Scripts: Groundtruth Generation at the Geusan Ulun Museum

Aug 2026 · Engineering, Technology & Applied Science Research · 0 citations · 13 references

TL;DR

The reproducible engineering process for building a ground truth-validated dataset for deteriorated Old Sundanese manuscripts from the Geusan Ulun Museum in Indonesia is described and can be applied to low-resource historical scripts, connecting cultural heritage and computational modeling.

Abstract

While cultural preservation and the digitization of historical manuscripts have a long history, low-resource writing systems continue to be poorly represented due to the absence of reliable ground truth datasets. This paper describes a reproducible engineering process for building a ground truth-validated dataset for deteriorated Old Sundanese manuscripts from the Geusan Ulun Museum in Indonesia. The Old Sundanese Manuscripts Dataset (OSMD) was constructed by combining controlled digitization, contrast enhancement, edge-based object extraction, morphological image processing, expert annotation, and optimized dataset packaging. The dataset contains 1,300 word-level samples for five consonant classes: Ha, Na, Ca, Ra, and Ka. To assess the potential of the dataset, baseline classification experiments were conducted using HOG+SVM and a lightweight CNN, where the latter outperformed (p<0.01) with 0.842 accuracy and 0.833 macro F1-score. The absence of contrast enhancement resulted in a notable drop in performance, and cross-validation demonstrated low variance (±0.006). The process followed in this study can be applied to low-resource historical scripts, connecting cultural heritage and computational modeling.

Read PDF

Similar papers

Conference Sep 2026

An AI-Assisted Classification-Driven Approach for Enhancement and Analysis of Palimpsest Readability

Palimpsests are ancient manuscripts in which the original text has been partially erased and overwritten by a later one, which makes them very difficult to decipher. This study examines the possibilities for analyzing, evaluating, and classifying the legibility of such manuscripts using modern image processing and mach...

G. Dimitrov, P. Tsvetkova, Pavel S. Petrov et al. · 0 citations
Open access Sep 2026

Formation of a Unified Alphabet of Ancient Turkic Dialects for Image Annotation

This paper is devoted to the development of a unified alphabet of Ancient Turkic dialects for solving the problems of automated recognition of Orkhon–Yenisei runic inscriptions. The relevance of the study is determined by the fragmentation of existing rune classifications and the absence of a unified system for correlatin...

A. D. Borodina, M. Kyzyl-ool, R. Kochkarov · 0 citations
Preprint Aug 2026

OCR-Based Field Extraction for Archaeological Pottery Metadata: The CENTURIA Dataset

Pottery is a primary source for reconstructing the chronological and economic dimensions of past societies. Archaeologists often document ceramic finds through technical drawings and handwritten metadata. This metadata is critical for dating, provenance attribution, and cross-site comparison, but remains inaccessible t...

Gissu Valentina Naghavi, Dominik Hagmann, M. Kampel et al. · 0 citations
Preprint Aug 2026

A Comparative Evaluation of Digitization Pipelines for Historiographical Sources

Purpose: The digitization of historical documents presents fundamental challenges for modern information retrieval and Artificial Intelligence (AI) systems. Optical character recognition (OCR) errors in source corpora propagate through retrieval-augmented generation (RAG) pipelines, compromising the factual accuracy of...

Marina Gómez Rey, Patricia Callejo, M. Muñoz-Organero et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.