Skip to content
Review Open access

Image-Based Fusion of Textual and Visual Information for Product Classification in Videos

2026 · IEEE Access · Vol 14, pp. 141173-141191 · 0 citations · 47 references

TL;DR

A hybrid multimodal framework combining intermediate cross-attention fusion with subsequent late fusion is proposed, which provides the strongest overall performance.

Abstract

With the rapid shift of shopping, marketing, and review content toward video-centric formats, automatic product-category recognition in videos has become increasingly important. However, existing image-based product classification research has largely focused on clearly visible products captured in controlled settings, which do not reflect the complexity of real-world video environments. Variations in lighting, viewpoint, and background make classification difficult, particularly for visually similar categories. To address these challenges, we propose a hybrid multimodal framework combining intermediate cross-attention fusion with subsequent late fusion. An object-detection branch identifies candidate products and extracts region-of-interest visual features, while Optical Character Recognition (OCR) extracts frame-level text that is processed by a fine-tuned multilingual Bidirectional Encoder Representations from Transformers (BERT) model. YOLOv9e is used as the object detector. Multi-head cross-attention integrates textual and visual representations before classification, and the resulting intermediate-fusion prediction is combined with visual and textual class-confidence outputs using an artificial neural network late-fusion classifier. Robust hyperparameters were selected using clean and noise-augmented validation data. Experiments on a self-collected cosmetic dataset containing 13,349 images, 17,875 annotated object instances, and 38 product categories show that the proposed framework achieves a micro F1 of 0.7520, a macro F1 of 0.7240, and a matched classification accuracy of 0.8462, outperforming the standalone YOLOv9e baseline by 10.44, 10.65, and 11.75 percentage points, respectively. Combining intermediate and late fusion provides the strongest overall performance.

Read PDF

Similar papers

Conference Open access Sep 2026

Auxiliary text-guided image restoration for image-text matching

A new ITM framework that improves the model's discriminative performance by focusing on localized core attributes and has a better robustness in handling highly similar hard negatives, which provides a new way to address cross-modal hard sample discrimination.

Kuang-Rong Hao · 0 citations
Review Sep 2026

A Survey on Text-Based Person Search.

This work provides a systematic review of TBPS by formalizing the problem setting and presenting a structured taxonomy of representative methods along four key technical dimensions: external knowledge-based approaches that leverage semantic priors to alleviate the inherent modality gap between visual and textual data.

Hai-Tao Shi, Meng Liu, Yong-Qi Li et al. · 0 citations
Open access Sep 2026

A multi-expert approach to content-based image retrieval using feature fusion and late re-ranking

As digital data rapidly grows, content-based image retrieval (CBIR) has become important for optimizing collections of visual data. This work proposes a retrieval framework which operates in two stages and improves accuracy by using systematic fusion of features. In the first stage, first-stage wide-scope descriptors c...

Ali Abdulazeez Mohammed Baqer Qazzaz, Yousif Samer Mudhafar · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.