Skip to content
Conference

Multi-scale local perception video topic recognition method based on semantic-guided feature pyramid

2026 · Poster Volume 0007 The 2026 Twenty-Second International Conference on Intelligent Computing July 23-26, 2026 Toronto, Canada · pp. 3571-3583 · 0 citations

TL;DR

A semantic-guided multi-scale feature pyramid learning method that significantly outperforms existing methods in scenarios requiring fine-grained local feature recognition, especially in topics such as "laboratory," "medical surgery," and "cooking process".

Abstract

Video topic recognition faces core challenges such as the limitations of static representations, cross-modal semantic misalignment, insufficient coverage of single-scale features, and weak temporal dynamic modeling. This paper finds that existing methods have significant deficiencies in local detail perception, leading to the neglect of key local cues such as "test tubes" in "laboratory" scenes, or the inability to simultaneously capture global environment and local details in "wedding" scenes. To address this, we propose a semantic-guided multi-scale feature pyramid learning method. The core innovation lies in the design of a temporal-semantic guided multi-scale local feature extractor (MLFE). This module can not only handle spatial multi-scale features but also incorporate temporal dynamics and textual semantic guidance to achieve adaptive scale selection. Based on this, we have constructed a complete recognition framework, including an improved cross-frame communication mechanism and a multi-granularity dynamic cue generation module. Experiments on benchmark datasets such as Kinetics-400 and Kinetics-600 show that our method significantly outperforms existing methods in scenarios requiring fine-grained local feature recognition, especially in topics such as "laboratory," "medical surgery," and "cooking process.

View source

Similar papers

Open access Aug 2026

Multi-scale global-local collaborative learning for accurate significant object detection

This work proposes MSGAN, a multi-scale global-local collaborative learning framework that integrates multi-scale mixed convolution and adaptive global-local attention to enhance feature representation and advances robust SOD for complex real-world scenarios and provides insights into attention-guided visual perception...

Jia-Yin Liu, Zetong Wang, Yu-Yuan Shen et al. · 0 citations
Open access Aug 2026

Action Recognition Method Based on Multi-Scale Dilated Feature Fusion and Decoupled Spatiotemporal Attention Pooling

Video action recognition requires the joint modeling of spatial appearance information and temporal dynamics. However, existing efficient action recognition methods based on two-dimensional convolution still have limitations in representing multi-scale spatial cues and aggregating key spatiotemporal information. To add...

Han-Bo Zhang, Jing Huang · 0 citations
Preprint Sep 2026

Text-Video Retrieval via Multi-Dimensional Saliency Assessment and Granularity-Aware Query Decomposition

Text-video retrieval, which aims to bridge visual and textual modalities by learning a joint embedding space, has become a crucial task in multimodal intelligence. Despite extensive efforts to mitigate visual redundancy, previous methods typically rely on a single-aspect criterion to assess visual importance, overlooki...

Shu-Quan Wei, Xi Chen, Xu Chen et al. · 0 citations
Open access Sep 2026

Pyramid-guided multi-scale self-attention and channel–spatial refinement for occlusion-robust face recognition

A pyramid-guided multi-scale attention framework based on scale alignment and reliability-aware feature refinement improves occlusion robustness without sacrificing clean-face recognition performance, indicating its practical potential for identity verification and access-control applications involving masks, glasses,...

Qi-Nan Zhu · 0 citations
Open access Sep 2026

A small and dense target detection framework based on feature perception and adaptive fusion

Small and dense object detection remains challenging in complex visual scenes. Repeated downsampling weakens discriminative features of tiny objects, while dense object distributions cause severe feature overlap and semantic ambiguity. To address these challenges, this paper proposes Enhanced Feature-Aware YOLO (EFA-YO...

Zhen Zhang, Xu Xie, Yi Zhang et al. · 0 citations
Open access Aug 2026

MAEF-Net: An Efficient Multi-Scale Attention-Enhanced Feature Fusion Network for Remote Sensing Object Detection

(1) Objective: Remote sensing object detection faces significant challenges, including complex background interference, large variations in target scales, and insufficient multi-scale feature representation, which often result in missed detections of small objects, inaccurate localization, and inadequate feature fusion...

Hongyan Shi, Xiaofeng Bai, Chenshuai Bai · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.