Skip to content
Conference

AART: image classification with adaptive attention region transformer

Aug 2026 · International Conference on Digital Image Processing · Vol 14351, pp. 143510Q - 143510Q-11 · 0 citations · 27 references
Engineering

TL;DR

The Adaptive Attention Region Transformer (AART) is proposed, which dynamically discriminates between regions based on their saliency, and achieves significant improvements over existing methods without requiring pre-training, validating its effectiveness in adaptive region processing.

Abstract

The Visual Transformer (ViT) has demonstrated powerful capabilities in modeling patch-wise attention for image classification. However, existing approaches typically treat all image regions uniformly, neglecting their inherent differences in importance. To address this limitation, we propose the Adaptive Attention Region Transformer (AART), which dynamically discriminates between regions based on their saliency. Our method begins by identifying key regions through density analysis of feature points, where the centroid of the densest cluster defines attention regions, with remaining areas designated as non-attention regions. We then implement differentiated feature extraction: small convolutional kernels capture fine-grained details from attention regions, while large kernels extract coarse-grained features from non-attention regions. This multi-scale feature extraction strategy enables more efficient representation learning. The resulting features are integrated and processed through Transformer blocks to learn comprehensive self-attentive representations. Extensive evaluations on CIFAR-10 and CIFAR-100 demonstrate that AART achieves significant improvements over existing methods without requiring pre-training, validating its effectiveness in adaptive region processing.

View source

Similar papers

Open access Sep 2026

A data efficient pyramid vision transformer for image classification

A novel data efficient pyramid vision transformer (DE-PVT), designed to train on limited datasets by utilizing a teacher-student approach and linear computational complexity relative to the number of patches, achieved through a linear spatial reduction mechanism is introduced.

Gazi Jannatul Ferdous, Medhi Hasan Chowdhury, Md. Azad Hossain et al. · 0 citations
Sep 2026

ALSRFormer: An adaptive transformer with dynamic window attention and multi-scale deformable feed-forward network for remote sensing image segmentation.

High-resolution remote sensing images present considerable challenges for semantic segmentation due to their complex object structures and extensive spatial distribution. Effective segmentation requires capturing fine-grained local details while simultaneously modeling long-range dependencies. Convolutional Neural Netw...

Zi-Qi Jia, Jia-Wei Zhang, Dong-En Guo et al. · 0 citations
Open access Aug 2026

Adaptive Feature Integration in CNN–Transformer Networks for Efficient and Interpretable Visual Classification

A novel adaptive fusion framework that adaptively combines CNN and Transformer features through learnable gating, attention-based feature integration, and explainable-AI methods is developed, intended to improve both computational efficiency and model interpretability.

Komal Sharma, Monika Sainger · 0 citations
Open access Sep 2026

An improved Inception Transformer with Bi-Level Routing Attention and multi-scale feature collaborative enhancement for Chinese ancient architecture image classification

Chinese ancient architecture embodies China’s cultural heritage but remains challenging to classify because of spatially dispersed key components, high inter-class similarity, and multi-scale visual characteristics. To address these challenges, we propose BDM-ViT, an improved Vision Transformer model based on the Inc...

Xin-Yang Wang, Yong-Sheng Zhang, Yu-Xun Peng et al. · 0 citations
Open access Sep 2026

Vision Transformer Architectures for Next-Generation Image Classification: Attention-Driven Visual Understanding

Vision Transformer (ViT) architectures have emerged as powerful alternatives to conventional convolutional neural networks for image classification because they model long-range visual dependencies through self-attention. This paper presents a software-based image classification framework that uses a pre-trained ViT-Ba...

Sadeqa and Dr. Bitla Prabhakar · 0 citations
Open access 2026

Efficient Hybrid Transformer-CNN Network for Real-World Image Denoising

Convolutional neural networks (CNNs) have long played a central role in computer vision due to their fast computation speed and high image feature extraction performance. CNNs can effectively learn diverse visual features ranging from low-level to high-level representations, resulting in high computational efficiency....

Ahhyun Lee, Sungkyun Shin, Dongsun Kim · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.