A novel data efficient pyramid vision transformer (DE-PVT), designed to train on limited datasets by utilizing a teacher-student approach and linear computational complexity relative to the number of patches, achieved through a linear spatial reduction mechanism is introduced.
Abstract
Transformer-based models such as the vision transformer (ViT) have seen widespread use in computer vision tasks. However, while ViT demonstrates excellent performance, it requires a massive dataset and the process of computing self-attention between patches has quadratic complexity. To handle these challenges, the vision transformer needs to be data-efficient to train effectively on smaller datasets and its computational complexity should scale linearly with the number of image patches. In response, this paper introduces a novel data efficient pyramid vision transformer (DE-PVT), designed to train on limited datasets by utilizing a teacher-student approach and linear computational complexity relative to the number of patches, achieved through a linear spatial reduction mechanism. In this research, patches are extracted from images using soft-splitting mechanism to capture local continuity and fine-grained details of images. The teacher model is a customized lightweight convolutional neural network that imparts knowledge to the transformer-based student model for classifying images. Moreover, a weighted loss function is employed to compute the overall cross-entropy loss in the proposed model. The final classification of the test image is determined by considering both the teacher and student models, with the contribution of each model being mathematically defined. To validate the model, the DE-PVT framework was trained on well-known benchmark image datasets including ImageNet-1K (32 × 32), CIFAR-10 and CIFAR-100, achieving F1-scores of 90.92%, 97.76%, and 94.59% respectively. These results suggest that the proposed model could positively impact computer vision tasks, particularly where data availability and rapid computation are critical concerns.
Convolutional neural networks (CNNs) have long played a central role in computer vision due to their fast computation speed and high image feature extraction performance. CNNs can effectively learn diverse visual features ranging from low-level to high-level representations, resulting in high computational efficiency....
Vision Transformer (ViT) architectures have emerged as powerful alternatives to conventional convolutional neural networks for image classification because they model long-range visual dependencies through self-attention. This paper presents a software-based image classification framework that uses a pre-trained ViT-Ba...
Sadeqa and Dr. Bitla Prabhakar· International Journal of Adv...· 0 citations
The Adaptive Attention Region Transformer (AART) is proposed, which dynamically discriminates between regions based on their saliency, and achieves significant improvements over existing methods without requiring pre-training, validating its effectiveness in adaptive region processing.
This study trains a hybrid deep learning system to efficiently recognize photos using the Caltech-256 dataset and findings validate the proposed hybrid design's provision of a strong and efficient model.
C. Patel· International Journal of Adv...· 0 citations
The ConvNeXt-Tiny architecture combined with Test-Time Augmentation is used in this paper to present a strong deep learning system for multi-class natural scene image classification, incorporating the design principles of Vision Transformers.
Mohammed Hamid Alkubaisi, Issa Mohammed Mishaal, B. T. Sabri et al.· 0 citations
A novel adaptive fusion framework that adaptively combines CNN and Transformer features through learnable gating, attention-based feature integration, and explainable-AI methods is developed, intended to improve both computational efficiency and model interpretability.
Komal Sharma, Monika Sainger· International journal of com...· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.