Skip to content
Preprint

Confidence-Aware Ensemble and Long-Word Refinement for Artistic Text Recognition

Aug 2026 · 0 citations · 28 references
Computer Science

TL;DR

A confidence-aware ensemble that combines SVTRv2, PARSeq, and MAERec after fine-tuning on the official training split is proposed, and an error analysis of all remaining mistakes shows that 48% are associated with labeling issues, visual ambiguity, or illegible samples, highlighting the value of diagnostic reporting for future ATR benchmarks and models.

Abstract

Artistic Text Recognition (ATR) remains challenging because word images often combine decorative fonts, curved layouts, object-like characters, clutter, and severe distortions. This paper studies WordArt-V1.5 as a standardized benchmark for this setting and evaluates recent scene and artistic text recognizers under a common protocol. We propose a confidence-aware ensemble that combines SVTRv2, PARSeq, and MAERec after fine-tuning on the official training split. The ensemble selects predictions using the minimum confidence over disagreement positions, emphasizing characters that separate competing hypotheses. For long words, where a single character error can invalidate the whole prediction, we add a targeted refinement stage based on Needleman-Wunsch alignment and lexicon-guided correction. On the WordArt-V1.5 Test B split, the proposed system reaches 89.90% Word Recognition Accuracy, improving the best individual fine-tuned model by 1.77 percentage points. The long-word refinement produces a modest global gain, but improves the targeted long-word subset by 2.72 percentage points. Finally, an error analysis of all remaining mistakes shows that 48.8% are associated with labeling issues, visual ambiguity, or illegible samples, highlighting the value of diagnostic reporting for future ATR benchmarks and models. Our source code is available at https://github.com/lucas-azdias/Artistic-Text-Recognition/.

View source

Similar papers

Preprint Sep 2026

Can Scene Text Recognition Read Rare Compositions?

No configuration the authors test improves both the compositional corner and aggregate accuracy, and the rare-input long tail thus points to architectural change rather than added capacity.

Gen-Pei Zhang · 0 citations
Preprint Sep 2026

A Free Lunch? Adapting PP-OCRv6 for Historical Text Recognition

This work adapts PP-OCRv6, a recent compact text recognizer without strong language modeling, for historical line recognition and compares it with a conventional CRNN across generalized pretraining, domain-specific training, corpus-level fine-tuning, and manuscript-specific few-shot adaptation on multilingual Latin- an...

Benjamin Kiessling · 0 citations
#computer vision Preprint Aug 2026

Towards a Joint Khmer Text Recognition and Word Segmentation

Experimental results show that the proposed model can not only recognize characters in document images but also locate word boundaries, removing the need for an extra word segmentation step in a conventional sequential pipeline.

Marry Kong, Rina Buoy, Sovisal Chenda et al. · 0 citations
Open access Aug 2026

BERT-based Automatic Error Correction System for Chinese Learners

Improved robustness and explainability for automatic Chinese learner error correction is demonstrated by an alignment consistency loss to ensure character-level consistency, and the combination of word-segmentation augmentation and multi-reference soft-label training to reduce conflicts caused by segmentation differenc...

Y.-L. Diao, W. Gao · 0 citations
#computer vision Preprint Sep 2026

On the Design Fundamentals of Pixel Text Representation Learning

This work investigates the fundamental design principles required for robust visual text representation learning and trains Pixel Linguist II, a native-resolution vision encoder trained with on-the-fly rendering, unified contrastive grounding, and a multilingual curriculum over 280M training examples.

Chaohao Yuan, Rui-Feng Yuan, Zhuoxu Huang et al. · 2 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.