DIMTrack: A Vehicle Multi-Object Tracking Method Integrating Spatial Attention and Disentangled Memory Learning
Abstract
Vehicle multi-object tracking (MOT) is a fundamental perception task in intelligent transportation systems, providing essential trajectory information for traffic monitoring, management, and autonomous driving applications. However, vehicle tracking in complex traffic environments remains challenging due to appearance similarity, frequent occlusions, viewpoint variations, and illumination changes. These factors may cause unstable appearance representations and unreliable data association, resulting in identity switches and fragmented trajectories. To address these issues, this paper proposes a vehicle MOT framework integrating spatial attention and disentangled memory learning. Specifically, a spatial attention mechanism is introduced to enhance discriminative regions, such as vehicle structures and fine-grained appearance details, improving feature robustness. Furthermore, an identity–variation disentangled learning strategy is designed to separate identity-related representations from variation factors, such as pose and illumination changes, reducing cross-frame feature drift. Meanwhile, an identity memory mechanism is developed to model historical identity features, enhancing trajectory continuity after occlusion. Experimental results on the UA-DETRAC and CityFlow datasets demonstrate the effectiveness and generalization capability of the proposed framework. On the UA-DETRAC dataset, DIMTrack achieves a HOTA score of 64.5% and an IDF1 score of 84.8%, while maintaining reliable identity association in dense traffic scenarios. Further evaluation on CityFlow shows that the proposed method generalizes well to diverse traffic environments, achieving an IDF1 score of 85.3% with fewer identity switches. These results confirm the robustness and transferability of DIMTrack for complex vehicle MOT scenarios.