BEV-Space Temporal Fusion for Multi-Modal Beam Prediction in Vehicular mmWave Systems
Abstract
Sensing-assisted beam prediction in millimeter-wave (mmWave) vehicular systems requires effective fusion of complementary multi-modal observations across sequential timesteps. Existing methods fuse globally pooled one-dimensional features, discarding cross-modal geometric correspondence before temporal aggregation. This paper proposes a BEV-space temporal fusion framework in which camera, LiDAR, radar, and GPS observations are first projected into a shared Bird's-Eye View (BEV) coordinate system at each timestep, and cross-modal fusion is performed in this geometrically grounded 2D domain. A temporal transformer then aggregates compact per-timestep descriptors derived from the fused BEV features, enabling trajectory-level reasoning with richer cross-modal spatial context than methods that pool each modality before interaction. A calibration-free camera-to-BEV transformation based on cross-attention avoids dependence on precise camera extrinsics, which are often unavailable in BS deployments. Experiments on DeepSense 6G (Scenarios 32–34) achieve 86.52% distance-based accuracy (DBA), outperforming the TransFuser baseline by 8.88 percentage points. Ablation studies confirm that BEV-space fusion before temporal pooling consistently outperforms 1D-fusion baselines.