DMF-Net: Dual-Modal Fusion Network for Beam Prediction in Low-Altitude mmWave Communication
Abstract
Driven by the booming low-altitude economy, UAV-to-Ground (U2G) millimeter-wave (mmWave) communications urgently demand low-latency, high-reliability beamforming. Conventional beam alignment based on channel feedback or exhaustive codebook search incurs prohibitive overhead and latency in dynamic scenarios, necessitating perception-aided communications. Existing single-modal schemes suffer from insufficient feature representation and poor robustness, failing to adapt to U2G mmWave channels' strong time-variation and high spatial selectivity. Multi-modal schemes with three or more modalities bring redundant information, excessive complexity and high cost, incompatible with UAV payload limits. In contrast, GPS-RGB fusion offers complementary temporal motion and spatial geometric features, with native UAV compatibility, minimal latency and light weight, ideal for low-altitude U2G scenarios. Hence, we propose a Dual-Modal Fusion Network (DMF-Net), an end-to-end beam prediction framework for U2G mmWave communications. Taking GPS and synchronized RGB images as inputs, it encodes temporal features via GRU, extracts spatial features via multiscale Cross Stage Partial Darknet (CSPDarknet), and predicts optimal beam index via lightweight MLP after cross-modal fusion. Experiments on DeepSense6G show our method achieves about 10% absolute top-1 accuracy improvement over the GPS single-modal baseline, with top-5 accuracy over 96%, and boosts convergence efficiency and generalization, validating dual-modal fusion's effectiveness for low-altitude mmWave beam prediction.