Skip to content
Preprint

ControlRadio: Prompt-Driven Controllable Diffusion for Cross-Modal Radio Map Generation

Aug 2026 · 0 citations
Computer Science

TL;DR

ControlRadio is presented, a controllable generative framework that produces radio maps from natural-language descriptions and environmental layouts, including building structures and transmitter locations, while reducing computation time by more than four orders of magnitude compared with conventional simulation-based methods.

Abstract

Radio maps describe how wireless signals propagate across space and are essential for wireless communication, sensing, and network planning. However, constructing accurate radio maps traditionally requires either dense measurements or computationally expensive physical simulations, which limits scalability and real-time deployment. Recent advances in generative artificial intelligence offer a promising alternative, but existing approaches lack fine-grained control and physical consistency when applied to real-world wireless environments. Here we present \textbf{ControlRadio}, a controllable generative framework that produces radio maps from natural-language descriptions and environmental layouts, including building structures and transmitter locations. Joint semantic and spatial conditioning enables interpretable, propagation-plausible generation, while a controlled latent prior and layout-aware conditioning improve stability and structural consistency. Extensive experiments demonstrate that ControlRadio achieves state-of-the-art accuracy and strong generalization across diverse urban scenarios, while reducing computation time by more than four orders of magnitude compared with conventional simulation-based methods. Such results suggest a new paradigm for scalable and controllable wireless environment modeling, with broad implications for next-generation communication systems and data-driven radio sensing.

View source

Similar papers

#artificial intelligence Preprint Aug 2026

Physics-Unrolled Neural Operator for Wireless Field Modeling

This work proposes Physics-Unrolled Hybrid Neural Operator (PU-HNO), a three-stage cascade that predicts high-fidelity indoor radio maps from low-fidelity ray-tracing outputs and scene priors by progressively capturing reflection, diffraction, and scattering effects, rather than treating radio maps as generic images.

Rafid Umayer Murshed, Saif Ur Rahman, Mingyue Tang et al. · 0 citations
Preprint Jul 2026

Map as a Prompt: Learning Multi-Modal Spatial-Signal Foundation Models for Cross-scenario Wireless Localization

Accurate and robust wireless localization is a critical enabler for emerging 5G/6G applications, including autonomous driving, extended reality, and smart manufacturing. Despite its importance, achieving precise localization across diverse environments remains challenging due to the complex nature of wireless signals and their sensitivity to environmental changes. Existing data-driven approaches often suffer from limited generalization capability, requiring extensive labeled data and struggling to adapt to new scenarios. To address these limitations, we propose SigMap, a multimodal foundation model that introduces two key innovations: (1) A cycle-adaptive masking strategy that dynamically adjusts masking patterns based on channel periodicity characteristics to learn robust wireless representations; (2) A novel"map-as-prompt"framework that integrates 3D geographic information through lightweight soft prompts for effective cross-scenario adaptation. Extensive experiments demonstrate that our model achieves state-of-the-art performance across multiple localization tasks while exhibiting strong zero-shot generalization in unseen environments, significantly outperforming both supervised and self-supervised baselines by considerable margins.

Yong Chu, Xun Zhou, Zenglin Xu et al. · 1 citation
2026

Rapid Generation of Channel Knowledge Map by Joint Physics and Conditional Diffusion Models

A channel knowledge map (CKM) provides location-specific channel priors and can reduce the overhead of real-time channel state information (CSI) acquisition for 6G environment-aware communications. In practice, CKM generation is often constrained by sparse and noisy measurements due to the high cost of wireless data collection. In this paper, we propose PDiff, a physics-informed conditional diffusion framework for CKM generation under sparse observations. Specifically, PDiff incorporates an analytical free-space propagation prior to capture the dominant distance-dependent attenuation trend, and combines it with environmental geometry, observation masks, and sparse observations as structured conditions. These conditional inputs guide the generation process with explicit propagation-aware, environmental, and measurement constraints. To improve inference efficiency, we further develop Prop-Cache, a training-free acceleration mechanism that reuses slowly varying intermediate features across denoising steps to reduce redundant computation during sampling. Experiments on RadioMapSeer demonstrate that PDiff outperforms a wide range of baseline methods for CKM generation.

Yu Chen, Jiao Chen, Jian Tang et al. · 0 citations
Preprint Jul 2026

RadioDiff-v2: Generative Angular Radio Maps for Multi-Beam Selection and Localization

Angular radio maps describe the received-power distribution over the angle of arrival and underpin beam selection and receiver localization in sixth-generation (6G) networks. Predicting the angular power spectrum (APS) from geometry is difficult, because the mapping is ill-posed in non-line-of-sight (NLOS) conditions and must generalize to unseen environments. Distortion-minimizing regressors return the conditional mean, which over-smooths the spectrum and erases the multipath structure that downstream tasks need. We cast the task as a perception-distortion problem and propose RadioDiff-v2, a dual-branch one-dimensional diffusion transformer trained with flow matching. It couples periodic angular encoding, adaptive layer-normalization conditioning, a Fourier angular mixer, and joint velocity and clean-signal heads. A per-metric estimator portfolio reads every deployment quantity from this single model, so that samples carry the distribution, the clean-signal head supplies a regression-grade point estimate, Bayes-optimal rules select beams, and the conditional likelihood localizes the receiver. We prove that a concentrated conditional yields a straight probability-flow trajectory that one step integrates exactly, identifying deterministic transport as the correct inductive bias. On a zero-shot test of 99 environments and one million links, RadioDiff-v2 leads every baseline on every metric, with a 0.39 dB Wasserstein-1 distance, per-bin error below the regression baseline, a 2.43 dB eight-beam NLOS sweep loss, and a 20.6-pixel localization error with four base stations. Code is available at https://github.com/UNIC-Lab/RadioDiff-v2.

Xiucheng Wang, Jun Huang, Nan Cheng · 0 citations
Jul 2026

RadioTrace: Transmitter-Aware Diffusion for Radio Map Estimation Without Deployment-Time Fine-Tuning

RadioTrace is proposed, a novel RM estimation framework without deployment-time fine-tuning that tightly integrates sparse RSS measurements with a frozen pre-trained diffusion prior and achieves competitive performance with state-of-the-art learning-based methods under random sampling, and maintains strong reconstruction quality under restricted-area sampling.

Liu Yang, Qiang Li, Zhuo Cao et al. · 0 citations
Conference Jul 2026

LSK-RM: A Physics-Guided Large Selective Kernel U-Net for Radio Map Reconstruction

Radio map construction aims to infer dense received-power fields from environmental layouts, sparse observations, and physical priors. While crucial for environment-aware wireless systems, it remains challenging in complex urban scenes with building-induced non-line-of-sight (NLOS) shadows and dynamic blockages. Non-iterative methods, such as interpolation techniques, RadioUNet, and RME-GAN, often struggle to accurately model these complex obstruction effects. Conversely, iterative generative methods tend to produce physically implausible hallucinations in shadowed or strongly obstructed areas. To address these issues, we propose LSK-RM, a physics-guided Large Selective Kernel U-Net for one-stage dense radio map reconstruction. Specifically, an LSK-based encoder-decoder is introduced to adaptively aggregate local shadow-boundary details and long-range attenuation context within a single forward pass. Furthermore, we develop a multi-source physical prior representation that fuses environmental geometry, sparse measurements, and fast ray-tracing visibility cues. To suppress physically implausible energy leakage, we design a novel logarithmic physics-guided objective combining pixel-wise supervision with Laplacian and obstacle-boundary consistency. Experiments on the RadioMapSeer dynamic blockage dataset demonstrate that LSK-RM outperforms representative baselines, including RadioUNet, RME-GAN, RMDM, and RadioFlow. Notably, it achieves higher accuracy across quantitative metrics such as NMSE, and exhibits significantly better modeling performance in diffraction transitions and shadow regions.

Zhengyan Liao, Weidong Zou, Chunlei Wang et al. · 0 citations