A bio-inspired ventral–dorsal fusion module for enhancing visual representations
Abstract
Deep neural networks have achieved remarkable success in visual recognition; however, learning stable and discriminative visual representations under diverse imaging conditions remains a persistent challenge. Existing approaches often rely on increasing model depth or complex attention mechanisms, which improve performance but tend to entangle complementary visual cues and limit representation robustness. To address this issue, we propose the Visual Dual-Stream Fusion Module (VDFM), a lightweight and biologically inspired component that simulates the cooperative processing of the human dorsal and ventral pathways. By integrating orientation-sensitive and color-sensitive cues through learnable gating, VDFM enhances the expressiveness and stability of visual feature representations. When integrated into ResNet- 50, VDFM consistently improves performance on Tiny-ImageNet, CIFAR-100, and STL-10 with minimal additional parameters and computational overhead. These results demonstrate that enhanced visual representations can be achieved through structured feature fusion rather than model expansion, providing an effective and biologically grounded approach for visual representation learning.