Skip to content
Preprint

Coupled Continuous-Discrete Generation for Scene Text Image Super-Resolution

Aug 2026 · 0 citations · 33 references
Computer Science

TL;DR

DualTSR is presented, a unified framework that formulates STISR as coupled continuous-discrete generation and achieves the best FID, LPIPS, ACC, and NED among the compared methods at both X2 and X4.

Abstract

Scene text image super-resolution (STISR) aims to recover visually plausible appearance while preserving character semantics from degraded inputs. Existing STISR systems often rely on externally generated priors or separate image and text models, resulting in error propagation and costly multi-stage inference. We present DualTSR, a unified framework that formulates STISR as coupled continuous-discrete generation. Conditional flow matching restores continuous image latents, while absorbing-state discrete diffusion reconstructs text tokens. Both processes share a multimodal transformer backbone, allowing the evolving image and text states to interact throughout generation without an external OCR prior at inference. On CTR-TSR, DualTSR achieves the best FID, LPIPS, ACC, and NED among the compared methods at both X2 and X4. On an aligned RealCE subset, it obtains the best FID, ACC, and NED with competitive LPIPS. Compared with DiffTSR at X4, DualTSR improves ACC by 12.78 percentage points while reducing the parameter count from 1.23B to 203M and end-to-end latency from 13.3s to 132ms. These results establish DualTSR as an accurate and efficient method for STISR.

View source

Similar papers

#machine learning Preprint Sep 2026

Panoptic Scene Program Diffusion Transformer

Modern text-to-image models produce high-fidelity images but still struggle with compositional prompts that require instance identity, attribute ownership, counting, spatial ordering, and role-sensitive relations. We introduce Panoptic Scene Program Diffusion Transformer (PSP-DiT), a diffusion-transformer architecture...

C. Maduabuchi · 0 citations
Preprint Aug 2026

PoseAdapter: Dual-Stream 2.5D Controllable Image Generation for Complex Multi-Object Scenes

PoseAdapter, a lightweight framework for high-fidelity 2.5D controllable image generation, and a Context-Aware Dual-Stream Representation, to resolve the generative trade-off between strict instance isolation and global coherence.

Yu-Feng Chi, Hui-Min Ma, Fan Gao et al. · 0 citations
Conference Open access Sep 2026

Auxiliary text-guided image restoration for image-text matching

Image-Text Matching (ITM) aims to establish deep semantic associations between visual content and textual descriptions. Existing methods usually have discrimination issues because of fine-grained semantic deviations, so it's hard to capture the complex correspondences between cross-modal entries. Only relying on alignm...

Kuang-Rong Hao · 0 citations
Aug 2026

GlyphTSR: Text-Rich Scene Image Super-Resolution Beyond Glyph Priors

Text-rich scene image super-resolution (TS-ISR) aims to recover high-quality images with legible text from degraded inputs, benefiting mobile photography and enhancing visual inputs for multimodal understanding. Existing methods rely on real-world image super-resolution or text image super-resolution, making it difficu...

Na Jiang, Yuxuan Qiu, Jia-Wei Zhang et al. · 0 citations
Preprint Sep 2026

RefCompose: Multi-Reference Image Generation via LoRA-Conditioned Diffusion

Filmmakers and visual artists routinely need to compose multiple references, actors, locations, props, cultural elements, into a single coherent shot, but existing tools either fail to scale past a handful of references or destroy fine grained subject identity in the process, since per reference tokenization scales mem...

Sai Sri Teja Kuppa, Parth Shinde, S. PriyadharsanBalaji et al. · 0 citations
Preprint Sep 2026

ENet-GP: Unified Document Image Restoration

Reliable document digitization in uncontrolled capture settings is challenging because real images exhibit multiple interacting degradations rather than a single isolated distortion. Documents thus captured are affected simultaneously by geometric distortions, like page warping, as well as photometric degradations such...

S. Burad, Aakanksha, A. Rajagopalan et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.