Review
Open access
2026
Image-Based Fusion of Textual and Visual Information for Product Classification in Videos
A hybrid multimodal framework combining intermediate cross-attention fusion with subsequent late fusion is proposed, which provides the strongest overall performance.
Yohwan Noh, Qi-Kang Deng, Dohoon Lee
· IEEE Access · 0 citations