DualTSR is presented, a unified framework that formulates STISR as coupled continuous-discrete generation and achieves the best FID, LPIPS, ACC, and NED among the compared methods at both X2 and X4.
Abstract
Scene text image super-resolution (STISR) aims to recover visually plausible appearance while preserving character semantics from degraded inputs. Existing STISR systems often rely on externally generated priors or separate image and text models, resulting in error propagation and costly multi-stage inference. We present DualTSR, a unified framework that formulates STISR as coupled continuous-discrete generation. Conditional flow matching restores continuous image latents, while absorbing-state discrete diffusion reconstructs text tokens. Both processes share a multimodal transformer backbone, allowing the evolving image and text states to interact throughout generation without an external OCR prior at inference. On CTR-TSR, DualTSR achieves the best FID, LPIPS, ACC, and NED among the compared methods at both X2 and X4. On an aligned RealCE subset, it obtains the best FID, ACC, and NED with competitive LPIPS. Compared with DiffTSR at X4, DualTSR improves ACC by 12.78 percentage points while reducing the parameter count from 1.23B to 203M and end-to-end latency from 13.3s to 132ms. These results establish DualTSR as an accurate and efficient method for STISR.
Modern text-to-image models produce high-fidelity images but still struggle with compositional prompts that require instance identity, attribute ownership, counting, spatial ordering, and role-sensitive relations. We introduce Panoptic Scene Program Diffusion Transformer (PSP-DiT), a diffusion-transformer architecture...
PoseAdapter, a lightweight framework for high-fidelity 2.5D controllable image generation, and a Context-Aware Dual-Stream Representation, to resolve the generative trade-off between strict instance isolation and global coherence.
Yu-Feng Chi, Hui-Min Ma, Fan Gao et al.· 0 citations
Image-Text Matching (ITM) aims to establish deep semantic associations between visual content and textual descriptions. Existing methods usually have discrimination issues because of fine-grained semantic deviations, so it's hard to capture the complex correspondences between cross-modal entries. Only relying on alignm...
Kuang-Rong Hao· International Conference on...· 0 citations
Text-rich scene image super-resolution (TS-ISR) aims to recover high-quality images with legible text from degraded inputs, benefiting mobile photography and enhancing visual inputs for multimodal understanding. Existing methods rely on real-world image super-resolution or text image super-resolution, making it difficu...
Na Jiang, Yuxuan Qiu, Jia-Wei Zhang et al.· IEEE Transactions on Image P...· 0 citations
Filmmakers and visual artists routinely need to compose multiple references, actors, locations, props, cultural elements, into a single coherent shot, but existing tools either fail to scale past a handful of references or destroy fine grained subject identity in the process, since per reference tokenization scales mem...
Sai Sri Teja Kuppa, Parth Shinde, S. PriyadharsanBalaji et al.· 0 citations
Reliable document digitization in uncontrolled capture settings is challenging because real images exhibit multiple interacting degradations rather than a single isolated distortion. Documents thus captured are affected simultaneously by geometric distortions, like page warping, as well as photometric degradations such...
S. Burad, Aakanksha, A. Rajagopalan et al.· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.