Skip to content
Preprint

A Free Lunch? Adapting PP-OCRv6 for Historical Text Recognition

Sep 2026 · 0 citations · 36 references
Computer Science

TL;DR

This work adapts PP-OCRv6, a recent compact text recognizer without strong language modeling, for historical line recognition and compares it with a conventional CRNN across generalized pretraining, domain-specific training, corpus-level fine-tuning, and manuscript-specific few-shot adaptation on multilingual Latin- and Arabic-script material.

Abstract

Despite impressive reported scores, large vision-language models have seen limited practical uptake in historical automatic text recognition because of their computational cost, dependence on large-scale pretraining, and hallucination. Historical ATR therefore continues to rely largely on compact CRNN line recognizers, which are visually grounded and trainable on modest data. Lightweight recurrence-free recognizers promise the accuracy of larger models with the practical advantages of CRNNs, yet have not been comprehensively evaluated on historical writing. We adapt PP-OCRv6, a recent compact text recognizer without strong language modeling, for historical line recognition and compare it with a conventional CRNN across generalized pretraining, domain-specific training, corpus-level fine-tuning, and manuscript-specific few-shot adaptation on multilingual Latin- and Arabic-script material. While PP-OCRv6 does not consistently outperform the baseline when trained from scratch, heterogeneous pretraining produces markedly better generalization. Comparisons with the Qwen3.5-based Medusa recognizer further show that fine-tuned PP-OCRv6 can outperform a large VLM tailored towards historical Latin-script HTR.

View source

Similar papers

Preprint Sep 2026

Exploring In-Context Learning for Handwritten Text Recognition

Handwritten Text Recognition (HTR) systems have become an indispensable tool for the digitization of historical documents. Not only do they cut down time and cost, but they also allow democratizing access and processing of their contents by generating their transcripts. However, literature in HTR currently focuses most...

Eric Ayllon, Abel Gandia, Jorge Calvo-Zaragoza · 0 citations
Preprint Aug 2026

Confidence-Aware Ensemble and Long-Word Refinement for Artistic Text Recognition

A confidence-aware ensemble that combines SVTRv2, PARSeq, and MAERec after fine-tuning on the official training split is proposed, and an error analysis of all remaining mistakes shows that 48% are associated with labeling issues, visual ambiguity, or illegible samples, highlighting the value of diagnostic reporting fo...

L. A. Dias, Henrique A. Schulz, Rafael Tadeu Machado de Miranda et al. · 0 citations
Preprint Aug 2026

SraVaani 1.0: Scaling Inclusive Speech Recognition for Indic Languages

SraVaani-1.0 achieves the lowest word error rate (WER) on a large number of language-dataset pairs while remaining competitive with the best-performing systems on high resource while being assessed exclusively on the VAANI benchmark.

Sujith Pulikodan, A. Basu, J. PavanKumar et al. · 1 citation · ⚡1
Preprint Sep 2026

Xiaomi-OCR-0 Technical Report

Compact OCR-specific vision-language models achieve strong document parsing performance, but often rely on costly supervision and focus primarily on visual-text reconstruction. We introduce Xiaomi-OCR-0, a unified 0.8B model for document parsing and OCR-centric understanding. We build an approximately 170M-sample OCR-c...

Xin Chen, An-An Du, Feng-Juan Feng et al. · 0 citations

HIT. A hybrid OCR methodology for document analysis for historical documents

This work provides a validated, privacy-preserving, and locally deployable solution for the high-fidelity transcription of sensitive human rights archives through a hybrid methodology that combines domain-specific fine-tuning for text recognition models with a novel anchoring mechanism to ground VLM generation.

Cristobal Sebastian Vasquez Rosel · 0 citations
Preprint Aug 2026

Cached LLM Probability Retrieval for Speech Recognition

Large language models (LLMs) enhance automatic speech recognition (ASR) by providing linguistic priors; however, their direct rescoring is costly because it requires evaluating every N-best hypothesis. This paper introduces"cached LLM probability retrieval,"which involves querying a local teacher LLM offline to obtain...

Sheng Li, Takahiro Shinozaki, Tatsuya Kawahara · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.