Jul 2026· 2026 8th International Conference on Electronics and Communication, Network and Computer Technology (ECNCT)· pp. 195-202· 0 citations· 24 references
Abstract
Radio map construction aims to infer dense received-power fields from environmental layouts, sparse observations, and physical priors. While crucial for environment-aware wireless systems, it remains challenging in complex urban scenes with building-induced non-line-of-sight (NLOS) shadows and dynamic blockages. Non-iterative methods, such as interpolation techniques, RadioUNet, and RME-GAN, often struggle to accurately model these complex obstruction effects. Conversely, iterative generative methods tend to produce physically implausible hallucinations in shadowed or strongly obstructed areas. To address these issues, we propose LSK-RM, a physics-guided Large Selective Kernel U-Net for one-stage dense radio map reconstruction. Specifically, an LSK-based encoder-decoder is introduced to adaptively aggregate local shadow-boundary details and long-range attenuation context within a single forward pass. Furthermore, we develop a multi-source physical prior representation that fuses environmental geometry, sparse measurements, and fast ray-tracing visibility cues. To suppress physically implausible energy leakage, we design a novel logarithmic physics-guided objective combining pixel-wise supervision with Laplacian and obstacle-boundary consistency. Experiments on the RadioMapSeer dynamic blockage dataset demonstrate that LSK-RM outperforms representative baselines, including RadioUNet, RME-GAN, RMDM, and RadioFlow. Notably, it achieves higher accuracy across quantitative metrics such as NMSE, and exhibits significantly better modeling performance in diffraction transitions and shadow regions.
Synchronized camera and wireless measurements observe the same scene through different physical channels. The central difficulty is that a representation learned in one deployment can fail when viewpoint, traffic, illumination, and propagation geometry change. This paper presents CM-MAE, a self-supervised vision--wireless pretraining framework for cross-scenario representation transfer. The evaluated real-data model uses only RGB frames and the measured 64-beam received-power vector available in DeepSense 6G; it does not use ray-traced paths, calibrated depth, or beam-index labels during pretraining. Its central pretraining term is a \emph{soft contrastive alignment loss}. Instead of making the synchronized image--wireless pair the only positive pair, this loss builds a target distribution from similarities between measured beam-power profiles, so nonidentical samples with similar directional responses are not forced apart as false negatives. A masked joint decoder provides the complementary local objective by reconstructing hidden visual patches and wireless angular clusters under modality dropout. After pretraining, a differential-rate fine-tuning rule lets a new fusion head adapt quickly while the encoders move slowly. Under a sequence-disjoint DeepSense 6G protocol, adding the soft alignment loss improves a matched linear-probe transfer average from 24.88\% to 29.49\%. Mild fusion fine-tuning reaches 77.38\% Top-1 accuracy on unseen Scenarios 6--8, and optional transductive normalization adaptation reaches 78.69\%. Since the fusion setting uses the contemporaneous 64-beam power vector at inference, these results should be read as representation-transfer diagnostics, not as proactive beam-prediction or reduced-sweeping claims.
RadioTrace is proposed, a novel RM estimation framework without deployment-time fine-tuning that tightly integrates sparse RSS measurements with a frozen pre-trained diffusion prior and achieves competitive performance with state-of-the-art learning-based methods under random sampling, and maintains strong reconstruction quality under restricted-area sampling.
Liu Yang, Qiang Li, Zhuo Cao et al.· IEEE Transactions on Wireles...· 0 citations
High-precision radio map construction is essential for emerging 6G Integrated Sensing and Communication (ISAC) applications, including digital twins and intelligent transportation. However, existing deep learning methods predominantly treat this as a pure image completion task, resulting in over-smoothed reconstructions that fundamentally erase high-frequency scattering signatures of dynamic physical entities such as hidden vehicles. To overcome this, we propose RadioVIL, an efficient two-stage framework that reformulates joint radio map inpainting and zero-shot vehicle localization as a prior-guided physical inverse problem. Specifically, we first train a Denoising Diffusion Probabilistic Model (DDPM) to capture the structural generative prior of the environment. During inference from highly sparse measurements, we employ a Diffusion-based Mediating Intermediate Layer Optimization (DMILO) algorithm. By optimizing an L1-regularized sparse deviation term, DMILO mathematically isolates vehicle scattering anomalies layer-by-layer without unfolding the entire denoising chain. Extensive experiments demonstrate that while conventional reconstruction baselines fail to detect hidden vehicles, and the zero-shot diffusion baseline achieves only limited detection ability due to forced semantic harmonization, RadioVIL preserves authentic physical textures, yielding the best LPIPS of 0.0587 in our evaluation. Uniquely, it unlocks accurate zero-shot vehicle localization directly from sparse radio maps, securing a 75.20% Recall and a 3.31-meter average error, paving a robust way for ISAC at the 6G edge.
Ruixin Zhao, Xiucheng Wang, Qiming Zhang et al.· 0 citations
This work introduces a physics- and tail-informed VAE-EVT (variational autoencoder-extreme value theory) framework that distinctly models both the bulk and tail distribution of SNR, and significantly outperforms the state-of-the-art GAN-based model.
A. Gamage, Niloofar Mehrnia, James Gross· 0 citations
The radio map characterizes the spatial distribution of spectrum resources within a region of interest and plays an important role in wireless network planning and spectrum management. In practice, observations are often sparse, making accurate radio map reconstruction challenging. Although deep learning–based methods can recover a radio map from sparse samples, they often suffer from two major limitations: a strong dependence on large amounts of training data and reconstructed results that may deviate from the actual physical distribution. To address this issue, this paper introduces dictionary factors to characterize radio propagation properties at different spatial scales and formulates radio map reconstruction as a multi-scale dictionary factor learning problem. Based on this formulation, we propose RadioMSDL-Net, a radio multi-scale dictionary learning network. The network consists of multiple layers with identical structures and progressively refines the radio map estimate in an iterative manner. In each layer, the dictionary factor update module learns propagation characteristics at different spatial scales, while the multi-level reconstruction module integrates cross-scale features to improve both the global structure and local details of the radio map. Extensive experiments demonstrate that RadioMSDL-Net consistently outperforms existing methods in reconstruction accuracy, computational efficiency, and cross-environment generalization.
Yazhou Sun, Longhui Wang, Xi Chen et al.· IEEE Transactions on Network...· 0 citations
Synthetic aperture radar (SAR) enables all-weather Earth observation; however, its inherent multiplicative speckle noise and geometry-dependent distortions pose significant challenges for SAR-to-optical image translation, often leading to structural deformation and degraded texture fidelity. To address these issues, this article presents Hierarchical MultiAxis Representation and Adaptive Residual Calibration Network (HARC-Net), an end-to-end Transformer-based regression framework that combines hierarchical multiaxis representation learning with statistics-guided skip-feature calibration to improve robustness and reconstruction quality. At the core of the proposed approach is a variable-axis sparse transformer (VASTormer) encoder, which integrates convolutional inductive bias with hierarchical multiaxis attention, including local block attention and sparse grid attention. This task-oriented encoder design enables efficient modeling of long-range dependencies while maintaining stable feature representations under speckle perturbations. To mitigate noise propagation in U-shaped architectures, we further introduce an adaptive dual attention and residual calibration (ADARC) module for skip connections. ADARC combines multistatistic spatial pooling (mean, max, min, and sum) with channelwise attention and learnable residual gating, effectively suppressing speckle-sensitive responses and improving semantic alignment between encoder and decoder features. Extensive experiments on two paired benchmarks, SEN1-2 and QXS-SAROPT, demonstrate that HARC-Net consistently achieves superior performance in both reconstruction quality and structural fidelity. The proposed method significantly reduces speckle-induced artifacts while preserving fine geometric details and linear structures. These results highlight the effectiveness of combining hierarchical local–global representation learning with statistics-guided feature calibration for robust cross-modal translation in remote sensing applications requiring geometrically consistent and noise-resilient optical reconstruction under adverse imaging conditions.
Yanghang Zhu, Zhijie Zhang, Mingsheng Huang et al.· IEEE Journal of Selected Top...· 0 citations