Skip to content
Open access

Large language models for OCR in cultural heritage: a comparative study on Slovene Folkloristic texts

Aug 2026 · International Journal on Document Analysis and Recognition · 0 citations · 16 references

Abstract

The digitization of historical and folkloristic texts presents significant challenges for optical character recognition (OCR), particularly when documents contain complex layouts, embedded illustrations, irregular typography, or non-standard language. This study provides a systematic evaluation of six OCR approaches on two Slovene-language heritage collections: typewritten folklore manuscripts with uniform formatting, and visually heterogeneous issues of the children’s magazine Ciciban . The methods compared include Tesseract, Tesseract with GPT 5.2 post-processing, GPT 5.2 direct transcription, LLaMA 4 Maverick, Nanonets OCR-3, and Qwen-VL-OCR. Performance was assessed using character error rate, word error rate, and complementary sequence-based metrics against manually aligned ground truth. Results indicate that direct multimodal and document-oriented systems achieve the strongest accuracy on typewritten texts, while performance on Ciciban is more sensitive to layout structure. These findings highlight the document-sensitivity of OCR performance and point to the need for adaptive, content-aware pipelines that dynamically integrate multiple OCR strategies. To our knowledge, the study provides the first systematic benchmark of LLM-based OCR for Slovene folkloristic materials, offering practical insights for cultural heritage digitization.

Read PDF