SFE-Net: semantic-guided parallel fusion for real-time semantic segmentation
Abstract
High-fidelity environmental perception is vital for autonomous systems under strict Size, Weight, and Power (SWaP) constraints. However, lightweight architectures often suffer from ”feature antagonism,” where high-level semantics inadvertently suppress critical spatial details during feature fusion. To address this, we propose the Semantic-gating Feature Enhancement Network (SFE-Net), which follows a decoupling-compensation-calibration paradigm. Specifically, a parallel hybrid neck segregates multi-scale features into dual streams, utilizing a detail enhancement module and a spatial compensation path to recover textures and long-range context. Furthermore, a semantic gating unit is introduced to calibrate heterogeneous signals and mitigate feature antagonism. Experimental results on Cityscapes demonstrate that SFE-Net achieves 74.33% mIoU with only 4.09M parameters and 2.16G FLOPs. The model reaches a real-time throughput of 140.99 FPS, significantly outperforming the SeaFormer baseline and providing an optimal balance between accuracy and efficiency for resource-constrained platforms.