Skip to content
Open access

A data efficient pyramid vision transformer for image classification

Sep 2026 · Discover Artificial Intelligence · Vol 6 · 0 citations · 50 references

TL;DR

A novel data efficient pyramid vision transformer (DE-PVT), designed to train on limited datasets by utilizing a teacher-student approach and linear computational complexity relative to the number of patches, achieved through a linear spatial reduction mechanism is introduced.

Abstract

Transformer-based models such as the vision transformer (ViT) have seen widespread use in computer vision tasks. However, while ViT demonstrates excellent performance, it requires a massive dataset and the process of computing self-attention between patches has quadratic complexity. To handle these challenges, the vision transformer needs to be data-efficient to train effectively on smaller datasets and its computational complexity should scale linearly with the number of image patches. In response, this paper introduces a novel data efficient pyramid vision transformer (DE-PVT), designed to train on limited datasets by utilizing a teacher-student approach and linear computational complexity relative to the number of patches, achieved through a linear spatial reduction mechanism. In this research, patches are extracted from images using soft-splitting mechanism to capture local continuity and fine-grained details of images. The teacher model is a customized lightweight convolutional neural network that imparts knowledge to the transformer-based student model for classifying images. Moreover, a weighted loss function is employed to compute the overall cross-entropy loss in the proposed model. The final classification of the test image is determined by considering both the teacher and student models, with the contribution of each model being mathematically defined. To validate the model, the DE-PVT framework was trained on well-known benchmark image datasets including ImageNet-1K (32 × 32), CIFAR-10 and CIFAR-100, achieving F1-scores of 90.92%, 97.76%, and 94.59% respectively. These results suggest that the proposed model could positively impact computer vision tasks, particularly where data availability and rapid computation are critical concerns.

Read PDF

Similar papers

Open access 2026

Efficient Hybrid Transformer-CNN Network for Real-World Image Denoising

Convolutional neural networks (CNNs) have long played a central role in computer vision due to their fast computation speed and high image feature extraction performance. CNNs can effectively learn diverse visual features ranging from low-level to high-level representations, resulting in high computational efficiency....

Ahhyun Lee, Sungkyun Shin, Dongsun Kim · 0 citations
Open access Sep 2026

Vision Transformer Architectures for Next-Generation Image Classification: Attention-Driven Visual Understanding

Vision Transformer (ViT) architectures have emerged as powerful alternatives to conventional convolutional neural networks for image classification because they model long-range visual dependencies through self-attention. This paper presents a software-based image classification framework that uses a pre-trained ViT-Ba...

Sadeqa and Dr. Bitla Prabhakar · 0 citations
Conference Aug 2026

AART: image classification with adaptive attention region transformer

The Adaptive Attention Region Transformer (AART) is proposed, which dynamically discriminates between regions based on their saliency, and achieves significant improvements over existing methods without requiring pre-training, validating its effectiveness in adaptive region processing.

Jing Liu, Xinyi Guo, Xin Zhang · 0 citations

ConvNeXt-Tiny for High-Accuracy Natural Scene Image Classification with Test-Time Augmentation

The ConvNeXt-Tiny architecture combined with Test-Time Augmentation is used in this paper to present a strong deep learning system for multi-class natural scene image classification, incorporating the design principles of Vision Transformers.

Mohammed Hamid Alkubaisi, Issa Mohammed Mishaal, B. T. Sabri et al. · 0 citations
Open access Aug 2026

Adaptive Feature Integration in CNN–Transformer Networks for Efficient and Interpretable Visual Classification

A novel adaptive fusion framework that adaptively combines CNN and Transformer features through learnable gating, attention-based feature integration, and explainable-AI methods is developed, intended to improve both computational efficiency and model interpretability.

Komal Sharma, Monika Sainger · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.