Backbone-Preserving Association Refinement for Visible-Thermal Multiple Object Tracking
Abstract
Visible-Thermal Multiple Object Tracking (VTMOT) benefits from complementary visible and thermal imagery, but identity preservation remains difficult under clutter, short occlusion, and temporary modality inconsistency. Recent progressive-fusion trackers such as PFTrack and MCTrack encode temporal and cross-modality cues in the backbone, yet their online association stage can still limit identity continuity. This paper presents a backbone-preserving association refinement for progressive-fusion VTMOT. The refinement combines score-ordered multistage matching with a hybrid association cost based on distance and Intersection over Union (IoU), while leaving the backbone, task head, and track lifecycle unchanged. On the VT-MOT benchmark, Association-Refined PFTrack (AR-PFTrack) improves the published PFTrack result by about 0.9 points in Higher Order Tracking Accuracy (HOTA) and 1.4 points in Identification F1 Score (IDF1), with a small localization-precision trade-off. Applying the same association configuration to MCTrack as an untuned transfer test improves HOTA and IDF1 but slightly lowers the Multiple Object Tracking Accuracy (MOTA), Multiple Object Tracking Precision (MOTP), and Detection Accuracy (DetA) scores. These results suggest that a lightweight association refinement can complement progressive-fusion representation learning at a small runtime cost and without added trainable parameters, although the gains are modest and scene-dependent.