Aug 2026· International Conference on Digital Image Processing· Vol 14351, pp. 143510Q - 143510Q-11· 0 citations· 27 references
Engineering
TL;DR
The Adaptive Attention Region Transformer (AART) is proposed, which dynamically discriminates between regions based on their saliency, and achieves significant improvements over existing methods without requiring pre-training, validating its effectiveness in adaptive region processing.
Abstract
The Visual Transformer (ViT) has demonstrated powerful capabilities in modeling patch-wise attention for image classification. However, existing approaches typically treat all image regions uniformly, neglecting their inherent differences in importance. To address this limitation, we propose the Adaptive Attention Region Transformer (AART), which dynamically discriminates between regions based on their saliency. Our method begins by identifying key regions through density analysis of feature points, where the centroid of the densest cluster defines attention regions, with remaining areas designated as non-attention regions. We then implement differentiated feature extraction: small convolutional kernels capture fine-grained details from attention regions, while large kernels extract coarse-grained features from non-attention regions. This multi-scale feature extraction strategy enables more efficient representation learning. The resulting features are integrated and processed through Transformer blocks to learn comprehensive self-attentive representations. Extensive evaluations on CIFAR-10 and CIFAR-100 demonstrate that AART achieves significant improvements over existing methods without requiring pre-training, validating its effectiveness in adaptive region processing.
A novel data efficient pyramid vision transformer (DE-PVT), designed to train on limited datasets by utilizing a teacher-student approach and linear computational complexity relative to the number of patches, achieved through a linear spatial reduction mechanism is introduced.
Gazi Jannatul Ferdous, Medhi Hasan Chowdhury, Md. Azad Hossain et al.· Discover Artificial Intellig...· 0 citations
High-resolution remote sensing images present considerable challenges for semantic segmentation due to their complex object structures and extensive spatial distribution. Effective segmentation requires capturing fine-grained local details while simultaneously modeling long-range dependencies. Convolutional Neural Netw...
A novel adaptive fusion framework that adaptively combines CNN and Transformer features through learnable gating, attention-based feature integration, and explainable-AI methods is developed, intended to improve both computational efficiency and model interpretability.
Komal Sharma, Monika Sainger· International journal of com...· 0 citations
Chinese ancient architecture embodies China’s cultural heritage but remains challenging to classify because of spatially dispersed key components, high inter-class similarity, and multi-scale visual characteristics. To address these challenges, we propose BDM-ViT, an improved Vision Transformer model based on the Inc...
Vision Transformer (ViT) architectures have emerged as powerful alternatives to conventional convolutional neural networks for image classification because they model long-range visual dependencies through self-attention. This paper presents a software-based image classification framework that uses a pre-trained ViT-Ba...
Sadeqa and Dr. Bitla Prabhakar· International Journal of Adv...· 0 citations
Convolutional neural networks (CNNs) have long played a central role in computer vision due to their fast computation speed and high image feature extraction performance. CNNs can effectively learn diverse visual features ranging from low-level to high-level representations, resulting in high computational efficiency....