Sep 2026· International Journal of Advanced Research in Science, Communication and Technology· 0 citations· 13 references
TL;DR
This study trains a hybrid deep learning system to efficiently recognize photos using the Caltech-256 dataset and findings validate the proposed hybrid design's provision of a strong and efficient model.
Abstract
The core computer vision problem of picture categorization has several domain-specific applications. This paper's goal is to talk about the significance of picture categorization in modern technology and society, as well as its ideas, techniques, and applications. Many computer vision systems rely on image classification, where images are automatically assigned to a predetermined category, based on the information they provide in the image. This study trains a hybrid deep learning system to efficiently recognize photos using the Caltech-256 dataset. The system makes use of both Bidirectional Long Short-Term Memory (BiLSTM) and Gated Recurrent Units (GRU). Dataset preprocessing includes standardization, label encoding, damaged image identification, and duplicate reduction. After separating the dataset into a train and test set, the following step is to extract features from each. The proposed GRU+BiLSTM model is evaluated in comparison to AlexNet, MobileNetV2, EfficientNet, and CNN models using ACC, PRE, REC, and F1-score (F1). Results for ACC (98.9%), PRE (99.1%), REC (99.3%), and F1 (100%) were better using the suggested method compared to the top deep learning models, according to the available experimental data. Findings validate the proposed hybrid design's provision of a strong and efficient
Vision Transformer (ViT) architectures have emerged as powerful alternatives to conventional convolutional neural networks for image classification because they model long-range visual dependencies through self-attention. This paper presents a software-based image classification framework that uses a pre-trained ViT-Ba...
Sadeqa and Dr. Bitla Prabhakar· International Journal of Adv...· 0 citations
A novel data efficient pyramid vision transformer (DE-PVT), designed to train on limited datasets by utilizing a teacher-student approach and linear computational complexity relative to the number of patches, achieved through a linear spatial reduction mechanism is introduced.
Gazi Jannatul Ferdous, Medhi Hasan Chowdhury, Md. Azad Hossain et al.· Discover Artificial Intellig...· 0 citations
A serial cascade of lightweight CNN and spectrum normalized GAN and spectrum normalized GAN, integrating CBAM attention mechanism is proposed, integrating CBAM attention mechanism, with good experimental results.
The ConvNeXt-Tiny architecture combined with Test-Time Augmentation is used in this paper to present a strong deep learning system for multi-class natural scene image classification, incorporating the design principles of Vision Transformers.
Mohammed Hamid Alkubaisi, Issa Mohammed Mishaal, B. T. Sabri et al.· 0 citations
Few-shot image classification remains difficult because a model must identify novel classes from only one or a few labeled examples while preserving discriminative local information. Metric-learning methods based on Earth Mover’s Distance (EMD) improve local correspondence by representing an image as a set of regional...
Huie Zhang, Mary Jane C. Samontet· International journal of com...· 0 citations
Deeper modern networks outperform the older AlexNet by a wide margin on CIFAR-10, and even a relatively compact ResNet can nearly match the accuracy of a much larger VGG16 in far less time.
Jinkuan Chen· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.