Skip to content
Open access

Vision Transformer Architectures for Next-Generation Image Classification: Attention-Driven Visual Understanding

Sep 2026 · International Journal of Advanced Research in Science, Communication and Technology · 0 citations · 1 references

Abstract

Vision Transformer (ViT) architectures have emerged as powerful alternatives to conventional convolutional neural networks for image classification because they model long-range visual dependencies through self-attention. This paper presents a software-based image classification framework that uses a pre-trained ViT-Base Patch16-224 model for single-image prediction and dataset-level evaluation. Images are collected from a Kaggle-style class-wise dataset and are processed through RGB conversion, resizing to 224 × 224 pixels, normalization, tensor conversion, and transformer-based inference. The proposed system also incorporates a graphical user interface that allows the user to select an image, visualize the input, and observe the predicted label and top-ranked class outputs. In addition, a performance evaluation module computes accuracy, precision, recall, F1-score, confusion matrix, and classification report. The uploaded result screenshots demonstrate successful qualitative predictions for a flower image, a teddy bear image, and a Labrador retriever image. The study shows that pre-trained Vision Transformer models can be effectively integrated into a flexible software-only classification framework suitable for academic demonstration, rapid prototyping, and future extension to custom image datasets.

Read PDF

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.