Reframing Historical Text Extraction: A Cross-Pathway Validation of OCR, LLM-Assisted Correction, and Direct Multimodal Transcription
Abstract
Historical document collections are increasingly available as digitised images and PDFs, but their conversion into reliable text remains affected by optical character recognition (OCR) errors, degraded pages, heterogeneous layouts, and domain-specific terminology. This study proposes a pathway-level framework for documenting and comparing conventional OCR, OCR followed by large language model (LLM)-assisted correction, and direct multimodal transcription. The framework is demonstrated using the Portuguese Agricultural and Forestry Surveys (1950–1958). A stratified validation sample of 45 pages was selected by visual quality, page type, and geographic coverage. Outputs were evaluated against manually verified reference transcriptions using content-normalised character error rate (CER) and word error rate (WER), document-condition analysis, paired statistical tests, and an entity-level semantic preservation assessment focused on place names, agricultural terms, and measurement expressions. Under the evaluated model and interface conditions, both LLM-based pathways produced lower mean CER and WER than the conventional OCR baseline, with the lowest values observed for direct multimodal transcription. Semantic preservation was also higher for the LLM-based pathways, although measurement expressions remained the most persistent risk, particularly in table-based pages. Downstream tasks were not directly evaluated. The findings support the framework as a method for validating text-extraction pathways before reuse, rather than establishing a universal ranking of tools.