Skip to content
Conference

FreqMamba: exposure-aware frequency decomposition with interleaved state scanning for HDR imaging

Aug 2026 · International Conference on Digital Image Processing · Vol 14351, pp. 1435107 - 1435107-15 · 0 citations · 29 references
Engineering

TL;DR

FreqMamba is proposed, a tri-branch frequency-aware framework built upon the synergy of CNNs, Swin Transformers, and Mamba, enabling joint exposure-spatial modeling with linear complexity in multi-exposure High Dynamic Range image reconstruction.

Abstract

Multi-exposure High Dynamic Range (HDR) image reconstruction aims to merge multiple Low Dynamic Range (LDR) images into a single HDR image with extended luminance range. Existing methods typically model spatial relationships and cross-exposure complementarity in isolation, overlooking their intrinsic coupling—where information reliability at each spatial location is fundamentally determined by its exposure level. Meanwhile, binary exposure masks widely used for region selection introduce boundary artifacts incompatible with the continuous variation of scene radiance, and frequency components are processed uniformly without accounting for their exposure-dependent signal quality. To address these limitations in a unified manner, we propose FreqMamba, a tri-branch frequency-aware framework built upon the synergy of CNNs, Swin Transformers, and Mamba. At its core, Multi-Exposure Interleaved State Scanning (MEISS) interleaves tokens from different exposures within quad-directional Mamba scanning, enabling joint exposure-spatial modeling with linear complexity. Luminance-Adaptive State Gating (LASG) further replaces hard threshold masks with a learnable soft gating mechanism that modulates state transitions based on multi-exposure luminance statistics, ensuring smooth radiance continuity. Building on this, Exposure-aware Frequency Decomposition and Recoupling (EFDR) decomposes features into frequency sub-bands processed by architecturally matched branches and fuses them via physics-guided reliability estimation under radiance consistency constraints. Extensive experiments on the Kalantari and Tel benchmarks demonstrate that FreqMamba achieves 42.43 dB PSNR-l and 0.9883 SSIM-l on challenging dynamic scenes, outperforming state-of-the-art methods with only 0.97M parameters and O(3N) linear complexity for cross-exposure modeling.

View source

Similar papers

2026

FD-HDRMamba: Frequency-Decoupled Mamba for Multi-Exposure HDR Reconstruction

Multi-exposure high dynamic range (HDR) imaging reconstructs scenes with large illumination variations by fusing multiple low dynamic range images, but large exposure gaps and scene motion often lead to ghosting artifacts, luminance inconsistency, detail degradation, and frequency imbalance. To address these challenges, we propose FD-HDRMamba, a frequency-decoupled HDR reconstruction framework that separately models global low-frequency structures and local high-frequency details. The proposed method first performs implicit feature-level alignment to reduce exposure and motion discrepancies, and then decomposes aligned features into frequency components. The low-frequency branch uses Mamba and a low-frequency-aware FFN to capture long-range dependencies and maintain global luminance consistency, while the high-frequency branch adopts residual feature distillation to enhance textures and structural details. Experiments on benchmark datasets show that FD-HDRMamba achieves competitive or superior performance in both quantitative metrics and visual quality, validating the effectiveness of frequency-decoupled HDR reconstruction.

Zhehan Gong, Wei Wang, Xiao Wang et al. · 0 citations
Open access Aug 2026

Dynamic Exposure-Adaptive Learning for Multi-Exposure Image Fusion Using RAW-Derived Training Pairs

Multi-exposure image fusion aims to expand the dynamic range of images by combining complementary information from differently exposed inputs. However, existing high dynamic range (HDR) reconstruction and exposure fusion methods often depend on HDR ground truth, tone mapping, or aligned multi-exposure data. To address these limitations, this study proposes a self-supervised high-exposure (HE) and low-exposure (LE) fusion framework for dynamic range expansion without requiring HDR ground truth. First, RAW-based exposure-range sampling augmentation generates diverse HE-LE training pairs from a single RAW image, expanding the training dataset while providing pixel-level aligned pairs without multi-shot acquisition. Second, a dynamic adaptive fusion module exploits complementary exposure information and balances local detail information with global structural information according to regional exposure characteristics. Fusion quality is progressively improved through Structural Similarity Index Measure (SSIM)-based coarse training and a stage-wise loss optimization strategy. Third, during inference, HDR-like enhancement and color compensation are applied to the LE image before fusion with the HE image to improve luminance, color consistency, and structural detail. Experimental results demonstrate that the proposed framework achieves stable dynamic range expansion without HDR ground truth and outperforms existing methods in structural preservation and visual quality, achieving the lowest BRISQUE (21.214), PIQE (33.979), and SSEQ (20.367) scores, as well as the highest MANIQA score (0.528), among all compared methods.

Seung Hwan Lee, Sung Hak Lee · 0 citations
Preprint Aug 2026

BC-IHV: Conditioning the Color Space for Stable Rectified-Flow Low-Light Enhancement

This work proposes Structure-Anchored Rectified Flow (SA-RF), which maintains correspondence through separate chromaticity/intensity stems, a scale-matched condition pyramid, and HybridAda, and introduces BC-IHV, a learnable Box--Cox polar color space whose analytically invertible intensity mapping controls the inverse-gradient dynamic range through a single exponent.

Yihao Ai, Zheng Chen, Yuanhao Cai et al. · 0 citations
Open access Jul 2026

Subband-Guided Hybrid Multi-Axis Attention Network for Frequency-Aware Image Super-Resolution

Single-image super-resolution (SISR) aims to reconstruct high-resolution (HR) images from low-resolution (LR) observations while preserving structural information and high-frequency detail. Although the Hybrid Multi-Axis Network (HMA) effectively combines local and nonlocal attention, its shallow input representation still mixes low-frequency structure, directional detail, and noise-like high-frequency components. This study investigates whether an explicit frequency prior can be introduced before the HMA backbone without substantially increasing computational cost. Two discrete-wavelet-transform front-ends are examined under the ×2 setting. HMA-WSB uses lightweight subband-specific processing and weighted fusion before a shared HMA backbone, whereas HMA-MSB introduces asymmetric multi-subband branches and cross-band fusion. Evaluation includes the reported external 15-image experiment, selected-image pilots from Set5, Set14, BSD100, and Urban100, and supplementary medical and texture-domain samples. The results show small, content-dependent differences rather than a consistent reconstruction advantage: the proposed variants are slightly favorable on several images containing dense multidirectional detail, but the original HMA remains stronger on other natural, medical, and periodic-texture samples. Computational analysis on an NVIDIA GeForce RTX 5070 with a 64×64 low-resolution input shows that HMA-WSB increases measured inference latency by 1.637% with negligible parameter and memory overhead. HMA-MSB increases latency by 5.078%, parameter count by 1.463%, and estimated FLOPs by 0.489%. These findings indicate that wavelet-guided subband processing is compatible with HMA and that WSB provides the more computationally economical extension. However, because the standard-dataset evaluation is based on selected images and a complete component-level ablation is not available, the results should be interpreted as preliminary evidence of a content-dependent quality-cost trade-off rather than proof of broad superiority.

Ching-Chun Chang, Tzu-Chuen Lu, Chin-Chen Chang · 0 citations
Open access Jul 2026

Trans2-CBCT: A Dual-Transformer Framework for Sparse-View CBCT Reconstruction

Cone-beam computed tomography (CBCT) with sparse projection views offers reduced radiation dose and faster scans but introduces severe streak artifacts and spatial coverage gaps. We address these challenges within a unified framework. First, we replace conventional UNet/ResNet encoders with TransUNet, a hybrid CNN–Transformer architecture that jointly models local details and long-range spatial context. It is adapted to CBCT reconstruction by concatenating multi-scale feature maps and introducing a lightweight attenuation-prediction head. Trans-CBCT outperforms the best baseline by 1.17 dB in PSNR and by 0.0163 in SSIM on LUNA16 with only six projection views. Second, we incorporate a neighbor-aware Point Transformer with explicit 3D positional encodings and a neighbor-aware attention module aggregating information from each point’s k-nearest spatial neighbors to enforce volumetric coherence. The resulting Trans2-CBCT achieves an additional 0.63 dB increase in PSNR and 0.0117 increase in SSIM over Trans-CBCT. In experiments with 6-10 views, Trans-CBCT and Trans2-CBCT consistently outperform all prior methods in both PSNR and SSIM on LUNA16. On the ToothFairy dataset, Trans2-CBCT leads in five of the six measurements, outperforming all baselines in PSNR. These results highlight the effectiveness of combining hybrid CNN–Transformer features with geometry-aware point-based reasoning for sparse-view CBCT reconstruction.

Minmin Yang, Yunhui Zhu, Huantao Ren et al. · 0 citations