Vision Transformer Architectures for Next-Generation Image Classification: Attention-Driven Visual Understanding
Abstract
Vision Transformer (ViT) architectures have emerged as powerful alternatives to conventional convolutional neural networks for image classification because they model long-range visual dependencies through self-attention. This paper presents a software-based image classification framework that uses a pre-trained ViT-Base Patch16-224 model for single-image prediction and dataset-level evaluation. Images are collected from a Kaggle-style class-wise dataset and are processed through RGB conversion, resizing to 224 × 224 pixels, normalization, tensor conversion, and transformer-based inference. The proposed system also incorporates a graphical user interface that allows the user to select an image, visualize the input, and observe the predicted label and top-ranked class outputs. In addition, a performance evaluation module computes accuracy, precision, recall, F1-score, confusion matrix, and classification report. The uploaded result screenshots demonstrate successful qualitative predictions for a flower image, a teddy bear image, and a Labrador retriever image. The study shows that pre-trained Vision Transformer models can be effectively integrated into a flexible software-only classification framework suitable for academic demonstration, rapid prototyping, and future extension to custom image datasets.