Multi-scale global-local collaborative learning for accurate significant object detection
Abstract
Salient Object Detection (SOD) remains a fundamental task in computer vision and visual computing, supporting applications ranging from image understanding to human-computer interaction. Existing methods still face two coupled challenges: insufficient modeling of multi-scale salient structures and imbalanced fusion between global semantic information and local details, which often lead to incomplete salient regions and blurred boundaries. To address these issues, this study proposes MSGAN, a multi-scale global-local collaborative learning framework that integrates multi-scale mixed convolution and adaptive global-local attention to enhance feature representation. Extensive experiments on the HKU-IS, ECSSD, PASCAL-S, and DUT-OMRON datasets demonstrate that our method achieves significant improvements in F-measure, MAE, and Em metrics, outperforming state-of-the-art approaches. Ablation studies validate the effectiveness of each core component. This work advances robust SOD for complex real-world scenarios and provides insights into attention-guided visual perception.