Image-Based Fusion of Textual and Visual Information for Product Classification in Videos
TL;DR
A hybrid multimodal framework combining intermediate cross-attention fusion with subsequent late fusion is proposed, which provides the strongest overall performance.
Abstract
With the rapid shift of shopping, marketing, and review content toward video-centric formats, automatic product-category recognition in videos has become increasingly important. However, existing image-based product classification research has largely focused on clearly visible products captured in controlled settings, which do not reflect the complexity of real-world video environments. Variations in lighting, viewpoint, and background make classification difficult, particularly for visually similar categories. To address these challenges, we propose a hybrid multimodal framework combining intermediate cross-attention fusion with subsequent late fusion. An object-detection branch identifies candidate products and extracts region-of-interest visual features, while Optical Character Recognition (OCR) extracts frame-level text that is processed by a fine-tuned multilingual Bidirectional Encoder Representations from Transformers (BERT) model. YOLOv9e is used as the object detector. Multi-head cross-attention integrates textual and visual representations before classification, and the resulting intermediate-fusion prediction is combined with visual and textual class-confidence outputs using an artificial neural network late-fusion classifier. Robust hyperparameters were selected using clean and noise-augmented validation data. Experiments on a self-collected cosmetic dataset containing 13,349 images, 17,875 annotated object instances, and 38 product categories show that the proposed framework achieves a micro F1 of 0.7520, a macro F1 of 0.7240, and a matched classification accuracy of 0.8462, outperforming the standalone YOLOv9e baseline by 10.44, 10.65, and 11.75 percentage points, respectively. Combining intermediate and late fusion provides the strongest overall performance.