CTF-Net for multimodal aided MIMO channel estimation
Abstract
Channel estimation is an essential element of MIMO-OFDM wireless communication systems. Conventional estimation techniques, including Least Squares (LS) and Minimum Mean Square Error (MMSE), alongside unimodal deep learning architectures such as Convolutional Neural Network (CNN), typically exhibit inadequate estimation precision under intricate terrain conditions. This paper introduces the CTF-Net, a cross-attention-guided multimodal fusion channel estimation network, to overcome the aforementioned challenges. CTF-Net establishes profound synergy between channel signal characteristics and terrain prior knowledge via CNN-based channel feature extraction, U-Net terrain encoding, and a bidirectional cross-attention fusion module, thus exceeding the performance constraints of conventional single-modal approaches. Experimental results from a real-terrain MIMO channel dataset developed through ray tracing indicate that the proposed CTF-Net attains a normalized mean square error (NMSE) as low as 0.0195 on the test set, signifying a reduction of 91.7% and 52.3% relative to conventional MMSE techniques and prevalent single-modal CNN-Attention models, respectively.