Skip to content

A Data-Model Jointly Driven Framework for Visible Light Positioning Using Harmonic-Enhanced Graph Neural Networks

Sep 2026 · IEEE Internet of Things Journal · Vol 13, pp. 40282-40298 · 0 citations · 33 references

Abstract

Visible light positioning (VLP), due to its widespread infrastructure deployment and high accuracy, has emerged as a highly promising key technology for Internet of Things (IoT). Current research mainly relies on either data-driven methods based on fingerprint features or model-driven methods based on geometric localization to estimate position. Although data-driven approaches can effectively cope with complex environmental disturbances, their performance is limited by insufficient exploitation of latent signal features on the one hand and strong dependence on training data on the other, resulting in limited generalization capability. In contrast, model-driven methods can adapt to different scenarios by leveraging physical models and geometric constraints, but they struggle to characterize complex interference patterns in dynamic environments. To address these issues, this article proposes a data-model jointly driven VLP framework that integrates the complementary strengths of both paradigms. First, at the framework level, a tightly coupled joint optimization scheme is constructed to integrate data-driven ranging with model-driven localization, preserving physical interpretability while leveraging the representation capability of deep learning. In contrast to conventional methods that discard harmonics as detrimental components, this article introduces a harmonic-enhanced ranging module that uses selected harmonic components as auxiliary structured spectral cues to improve the robustness of VLP ranging. The fundamental received signal strength (RSS) and selected harmonic RSS values of the modulation signal are incorporated into the ranging process, and a graph neural network (GNN)-based data-driven ranging model is developed to capture the structured relationships between the fundamental component and its harmonics, thereby improving ranging robustness in complex environments. Finally, to enable end-to-end optimization of the joint framework, a geometry-aware loss function is designed, allowing the model to jointly consider data fitting and physical geometric constraints during training, thereby coupling learned signal representations with model-driven geometric constraints.

View source

Similar papers

Preprint Jul 2026

Map as a Prompt: Learning Multi-Modal Spatial-Signal Foundation Models for Cross-scenario Wireless Localization

Accurate and robust wireless localization is a critical enabler for emerging 5G/6G applications, including autonomous driving, extended reality, and smart manufacturing. Despite its importance, achieving precise localization across diverse environments remains challenging due to the complex nature of wireless signals and their sensitivity to environmental changes. Existing data-driven approaches often suffer from limited generalization capability, requiring extensive labeled data and struggling to adapt to new scenarios. To address these limitations, we propose SigMap, a multimodal foundation model that introduces two key innovations: (1) A cycle-adaptive masking strategy that dynamically adjusts masking patterns based on channel periodicity characteristics to learn robust wireless representations; (2) A novel"map-as-prompt"framework that integrates 3D geographic information through lightweight soft prompts for effective cross-scenario adaptation. Extensive experiments demonstrate that our model achieves state-of-the-art performance across multiple localization tasks while exhibiting strong zero-shot generalization in unseen environments, significantly outperforming both supervised and self-supervised baselines by considerable margins.

Yong Chu, Xun Zhou, Zenglin Xu et al. · 1 citation
Preprint Aug 2026

Electromagnetic World Model for 6G: A Unified Framework for Joint Environment Reconstruction and Channel Prediction

The integration of sensing, communication, and intelligence is becoming a key enabler for sixth generation (6G) wireless systems, where intelligent terminals are expected to simultaneously support efficient link establishment and reliable environmental sensing. However, existing studies mainly exploit sensing information or communication information to address a single task, such as channel prediction or environment reconstruction. Motivated by the shared dependence of optical and radio-frequency signals on the surrounding environment, we propose the electromagnetic world model (EMWM), the first unified framework for joint environment reconstruction and channel prediction. EMWM learns a common electromagnetic representation with the potential to provide a modeling foundation for 6G tasks. Specifically, partial channel state information (CSI) and multi-view red-green-blue (RGB) images are encoded into CSI and visual tokens and jointly processed by a hierarchical world-model backbone with local and global aggregation. Based on the learned representation, a mixture-of-experts (MoE)-based CSI prediction head reconstructs the complete CSI, while a depth prediction head estimates multi-view depth maps that are further converted into three-dimensional (3D) point clouds. Moreover, a large-scale multi-modal dataset is constructed based on a campus digital twin. Experimental results show that EMWM outperforms conventional neural network and large language model (LLM) baselines in both CSI prediction and environment reconstruction, achieving a squared generalized cosine similarity (SGCS) of 0.9699 for CSI prediction while demonstrating robustness across different signal-to-noise ratio (SNR) conditions and zero-shot generalization at 28 GHz.

Yizhu Zhao, Li Yu, Jianhua Zhang et al. · 0 citations
Open access Jul 2026

Sum Rate Optimisation for IRS-Assisted VLC Systems with Time Delay Considerations Using Deep Q-Learning

Visible light communication (VLC) is a promising solution for high-speed indoor wireless connectivity, offering advantages such as license-free spectrum and enhanced physical-layer security. However, VLC performance is highly dependent on line-of-sight (LoS) conditions and is vulnerable to signal degradation caused by device orientation and dynamic obstructions. To address these challenges, this paper proposes a deep Q-Learning (DQL) framework for resource optimisation in intelligent reflecting surfaces (IRS)-assisted VLC systems. A realistic system model is developed that incorporates both LoS and non-line-of-sight (NLoS) com-ponents, while explicitly modeling frequency-domain effects due to IRS-induced time delay. The optimisation problem jointly considers IRS element allocation and user-LED association, aiming to maximise the system sum rate under practical constraints such as quality of service and IRS-induced time delay. A DQL algorithm is designed to learn efficient allocation strategies in high-dimensional dynamic environments. Simulation results show that the proposed DQL approach closely matches the achievable rate predicted by analytical models. Furthermore, the results highlight the critical impact of accounting for time delay and user orienta-tion when designing IRS-assisted VLC systems. The findings support the viability of learning-based, delay-aware optimisation for next-generation intelligent indoor VLC networks.

A. Hussen, Rashid Iqbal, A. Zoha et al. · 0 citations
Open access 2026

A Causal Temporal Convolutional Network for End-to-End Pose Estimation in Visible Light Positioning System

Visible Light Positioning (VLP) has emerged as a promising solution for high-precision indoor localization due to its immunity to electromagnetic interference, high spatial resolution, and integration with existing lighting infrastructure. However, conventional Received Signal Strength (RSS)-based localization approaches suffer from severe performance degradation under nonlinear optical channel conditions, measurement noise, and orientation-dependent signal variations. This paper proposes a novel end-to-end causal Temporal Convolutional Network (TCN) framework for simultaneous three-dimensional position and single-axis orientation (azimuth) estimation for receivers moving on a fixed horizontal plane in indoor VLP systems. Unlike conventional Extended Kalman Filter (EKF)-based localization methods that rely on first-order linearization of nonlinear Lambertian channel models, the proposed TCN exploits causal dilated convolutions to capture temporal RSS dynamics and nonlinear mobility patterns directly from sequential measurements. A soft-attention temporal pooling mechanism is further incorporated to suppress noisy and unreliable RSS observations. The proposed framework is evaluated in a realistic simulated indoor VLP environment with 180,000 training samples generated under additive white Gaussian noise, background illumination interference, and optical crosstalk conditions. Simulation results demonstrate that the proposed TCN framework significantly outperforms the conventional EKF approach in terms of convergence speed, positioning accuracy, and tracking stability. The proposed method achieves convergence within approximately 2-3 iterations, whereas the EKF requires nearly 8-10 iterations under identical conditions. Furthermore, the average convergence time is reduced by approximately 59% while maintaining stable steady-state estimation performance. Experimental results also show lower position and orientation estimation errors, reduced localization outliers, and improved trajectory tracking accuracy during dynamic circular motion scenarios. The proposed TCN-based VLP framework provides a computationally efficient and robust solution for practical next-generation indoor localization applications.

Sunita Khichar, Sushank Chaudhary, Amir Parnianifard et al. · 0 citations
Nov 2026

Active Meta-Learning for Few-Shot QoT Estimation in Optical Networks

Accurate quality of transmission (QoT) estimation, particularly generalized signal-to-noise ratio (GSNR) prediction, in newly deployed C+L-band optical networks is hindered by data scarcity and incomplete physical link information, leading to a cold-start problem for conventional deep learning methods. To address this issue, we propose PAL-MISA, a data-efficient framework combining Parameterized Active Learning (PAL) and Meta-Initialized Sparse Adaptation (MISA). PAL learns to select high-value measurements, while MISA provides a transferable initialization for rapid sparse adaptation to unseen physical environments. Validated on three representative network topologies, PAL-MISA accelerates convergence and reduces performance fluctuations compared with conventional training from scratch in the target domain. To achieve the same prediction error, PAL-MISA reduces the required measurement data volume by 80% and ultimately achieves a minimum mean absolute error (MAE) of 0.04 dB, offering a robust solution for digital-twin deployment in uncharacterized optical networks.

Haoyu Wang, Ruiyang Xia, Zanshan Zhao et al. · 0 citations
Open access Jul 2026

Computationally Efficient Deep Learning Approach Using IQ-MobNet for Radar DoA Estimation in Limited Snapshot Conditions

This paper presents a computationally efficient deep learning framework for accurate direction-of-arrival (DoA) estimation in portable radar applications. Leveraging a MobileNet architecture, the proposed model directly processes raw in-phase and quadrature-phase (IQ) data, enabling more effective learning of both spatial and temporal features. This direct input approach enhances DoA estimation accuracy, particularly under challenging conditions such as low signal-to-noise ratio (SNR) and limited snapshot scenarios. A unified training strategy is adopted for both single-source and multi-source target detection, ensuring consistency and robustness. Comprehensive simulation experiments demonstrate the proposed model’s competitive and robust performance across various conditions, including different SNR levels, closely spaced targets, and random off-grid angles. It also shows that our method achieves performance comparable to or better than recent deep learning approaches in several challenging scenarios, establishing its potential for resource-constrained environments where only low snapshot data are available. The proposed IQ-MobNet DoA estimation model achieves this competitive performance with substantially lower computational complexity, requiring only 0.24 million parameters and 0.42 million Floating Point Operations (FLOPs), representing a reduction of over 96% compared to the recent neural network models. To ensure practical applicability, the proposed IQ-MobNet framework is validated using real-world measured data, confirming its robustness beyond simulated environments.

Neeraja P. Kovilakam, Bindiya T. Sambasivan, Raghu C. Variyam · 0 citations

Related blog posts