A robust attention-guided multi-dimensional feature fusion network (LFAMF) for LF angular reconstruction that significantly outperforms state-of-the-art methods, particularly in maintaining structural integrity at occlusion boundaries and highly textured areas while ensuring superior angular consistency.
Abstract
Angular super-resolution (ASR) is a fundamental task in light field (LF) imaging, aimed at reconstructing a dense LF from sparsely sampled views. Despite significant progress, current methods often struggle to preserve consistency in complex scenarios such as severe occlusions and large-disparity regions. In this paper, we propose a robust attention-guided multi-dimensional feature fusion network (LFAMF) for LF angular reconstruction. The proposed framework comprises two synergistic stages: a multi-dimensional feature fusion stage and an attention-guided refinement stage. Specifically, we design a multi-stream subnetwork (MFNet) to extract intrinsic physical characteristics across the spatial, angular, EPI, and pseudo-video sequence domains. Simultaneously, a geometry-prior-based subnetwork (GSPNet) is incorporated to leverage scene structure for improved texture preservation. To effectively integrate these complementary streams, an attention-guided fusion subnetwork (AFNet) is employed to adaptively merge intermediate results. Extensive experiments on both synthetic and real-world datasets demonstrate that the LFAMF model significantly outperforms state-of-the-art methods, particularly in maintaining structural integrity at occlusion boundaries and highly textured areas while ensuring superior angular consistency.
Complex light field reconstruction requires recovering high-resolution light field data with spatial clarity, angular consistency, and occlusion boundary stability under limited viewpoint sampling conditions. This is a key issue in computational imaging, free-viewpoint display, 3D perception, and immersive interaction. Existing depth reconstruction methods can usually improve detail representation by utilizing complementary information between adjacent viewpoints. However, when sparse viewpoints, large parallax, weak texture, non-uniform noise, and occlusion boundaries coexist, viewpoint drift, edge blurring, and high-frequency texture loss are still prone to occur. To address these issues, this paper proposes a multi-scale feature fusion and adaptive optimization method, MSAF-AO, for complex light field reconstruction. This method first constructs a unified spatial-angle input representation, mapping sparse sub-aperture images to a shared feature domain. Then, it uses multi-scale feature branches to capture local texture, edge contours, and global parallax structure respectively. Furthermore, it designs adaptive fusion weights for occlusion awareness, enabling features of different scales to dynamically participate in reconstruction according to regional complexity. Finally, it constructs a joint objective function composed of reconstruction error, angular consistency, edge preservation, and fusion regularization. Experimental results show that MSAF-AO achieves higher PSNR, SSIM, and lower angular consistency error under complex parallax and noise conditions, and has more stable structural recovery capability in occluded areas.
Yudi Feng, Yaping Wang· International Conference on...· 0 citations
Most existing low-light image enhancement methods mainly rely on feature modeling in a single domain, making it difficult to simultaneously achieve global illumination correction, local detail restoration, and noise suppression. To overcome this limitation, we propose a low-light image enhancement network(MPNet) that leverages multi-domain prior attention, and introduce MPGSA (Multi-domain Prior Guided Self-Attention) as its core module. Specifically, MPGSA incorporates priors from the spatial, Fourier, and wavelet domains into the attention mechanism. Through complementary multi-domain modeling, it improves global illumination restoration, local texture reconstruction, and noise suppression. Extensive experimental results demonstrate that the proposed method not only achieves outstanding visual quality and competitive quantitative metrics on multiple public datasets, but also performs excellently on downstream tasks, showcasing strong generalization ability and broad application potential.
Shifan Yang, Qizhao Lin, Weilin Wu et al.· IEEE Signal Processing Lette...· 0 citations
Low-light image enhancement aims to improve visual visibility and perceptual quality under challenging illumination conditions. However, conventional convolutional neural networks (CNNs) are inherently limited in modeling long-range dependencies due to their restricted receptive fields, which often leads to insufficient global context modeling and suboptimal restoration results. To address this limitation, we propose MSHCDI-Net, a Multi-Scale Hybrid Cross-Domain Interaction Network that effectively integrates CNN and Transformer branches to jointly capture local texture details and global contextual relationships. Specifically, the proposed framework adopts a hierarchical encoder–decoder architecture to perform multi-scale feature extraction and progressive reconstruction. A cross-domain interaction mechanism is introduced to facilitate effective information exchange between convolutional and Transformer representations across multiple resolutions, enabling complementary modeling of fine-grained structures and long-range dependencies. Through adaptive feature fusion and multi-scale guidance, the network achieves improved structural consistency and detail restoration in low-light scenes. Extensive experiments on several public benchmarks demonstrate the effectiveness of the proposed method. MSHCDI-Net achieves 23.45 dB PSNR / 0.848 SSIM on LOL-v1, 23.74 dB / 0.910 SSIM on LOL-v2-synthetic, and 22.24 dB / 0.868 SSIM on LOL-v2-real, demonstrating competitive performance in both quantitative metrics and visual quality.
Bin Chen, Peitao Li, Chaobing Zheng et al.· PLoS ONE· 0 citations
Highlights What are the main findings? Proposes Structure–Detail Constrained Fusion (SDC-Fusion), a frequency-decoupled infrared and visible image fusion framework that separately models low-frequency structure and high-frequency detail. Restricts Rectified Flow to the high-frequency wavelet subbands and learns a gated few-step residual compensation trajectory from the visible high-frequency component to a local-directional-energy-guided target, rather than generating the complete fused image or latent representation. Develops a low-frequency structure preservation module that integrates a four-directional Mamba scan with multi-dilated depthwise convolutions, simultaneously maintaining global luminance integrity and optimizing local grayscale transitions. Ranks first on five of seven metrics and second on the remaining two metrics on both the MSRS and M3FD datasets, while using 0.535 M parameters and 67.5 G FLOPs. Abstract Existing end-to-end infrared–visible fusion methods often blur edges, smooth textures and weaken target-to-background contrast. We therefore propose Structure–Detail Constrained Fusion (SDC-Fusion), a frequency-decoupled framework with separate constraints on structure and detail. The proposed method employs the Haar wavelet transform to decompose the source images into low-frequency structural and high-frequency detail components. The high-frequency branch uses a local directional-energy prior to construct the target guidance. Unlike existing Rectified Flow-based approaches that operate on the full image or a generic latent representation, our gated module applies Rectified Flow only to the high-frequency wavelet subbands. It learns a few-step residual trajectory from the visible high-frequency coefficients to the target representation, enhancing infrared target boundaries and visible textures without altering low-frequency structure. In the low-frequency branch, adaptive weighting, four-directional Mamba scanning, and multi-dilation depthwise convolutions are integrated to preserve global luminance and background structure while optimizing local grayscale transitions. Comparative experiments against eleven representative fusion methods on the MSRS and M3FD datasets show that SDC-Fusion ranks first in SSIM, VIF, Qabf, SF, and PSNR on MSRS, and first in SSIM, VIF, Qabf, SD, and PSNR on M3FD, while ranking second in the remaining two metrics on each dataset. Relative to the strongest competing result, the largest improvements reach 10.44% in SF on MSRS and 5.72% in VIF on M3FD. The model contains 0.535 M parameters and requires 67.5 G FLOPs.
Xiaoxia Wang, Shuang Guo, Fengbao Yang et al.· Italian National Conference...· 0 citations
Depth super-resolution reconstructs high-resolution (HR) depth maps from low-resolution (LR) inputs with the aid of HR RGB guidance, but RGB edges often do not coincide with true depth discontinuities, causing texture copying and degraded geometric consistency. To address this problem, we propose Frequency–Geometry-Guided Network (FGGNet), a spatial–frequency fusion framework for RGB-guided depth map super-resolution. FGGNet introduces Multi-branch RGB-guided Convolution (MRGConv) to enhance RGB structural representations, a Geometry Prior-guided Fusion Module (GPFM) to filter geometrically inconsistent RGB responses using depth-derived priors, and radial complex spectral loss (RCSL) to emphasize boundary-related high-frequency components in the complex spectral domain. Experiments on NYU v2, Middlebury, Lu, and RGB-D-D show that FGGNet achieves competitive or superior reconstruction accuracy under synthetic and real-world degradation settings. Under the ×16 setting, FGGNet reduces RMSE by 13.7%, 22.8%, 18.5%, and 11.4% on the four datasets, respectively, compared with the average RMSE of five representative state-of-the-art methods. These results validate the effectiveness of combining geometric prior filtering with frequency-domain supervision for reliable depth reconstruction.
Zhiqiang Feng, Chong Zhang· Italian National Conference...· 0 citations