Tracker-Conditioned Diffusion Proposals for Multi-Object Tracking
Abstract
Diffusion-based detectors begin inference from noisy boxes, whereas tracking-by-detection pipelines usually localize each video frame independently. This study examines whether propagated tracker boxes can guide a frozen DiffusionDet detector without retraining. Confirmed tracks are propagated using their latest observation or a Kalman prediction, mapped to the detector’s latent box space, corrupted at a selected diffusion timestep, and mixed with random proposals for new-object discovery. On a KITTI development split, four low-noise proposals around each Kalman prediction improve mean car/pedestrian multiple object tracking accuracy (MOTA) by 1.29 points over a matched short-timestep random control and 0.66 points over full random initialization. Higher order tracking accuracy (HOTA) is effectively tied, while the identity F1 score (IDF1) is lower than with full random initialization. Transferred without retuning to MOT17 half-validation, the configuration averages 53.68 HOTA, 55.88 MOTA, and 63.91 IDF1 across three inference seeds. Mean HOTA improves by 0.99 points over full random initialization, primarily through higher recall at the cost of lower precision. The findings support a conditional detection-recall benefit rather than a consistent improvement in identity association.