Low-Frame-Rate Multi-Object Tracking (LFR-MOT) is proposed, a purely appearance-based tracker that removes motion prediction entirely and relies on re-identification (ReID)-based appearance features with a two-stage matching strategy to handle detection uncertainty.
Abstract
Surveillance and remote monitoring systems operating under bandwidth and storage constraints commonly record at extremely low frame rates, often as low as 1 frame per second (fps). At this temporal resolution, the core assumptions underlying conventional multi-object tracking (MOT) break down. Motion prediction based on the Kalman filter becomes unreliable because interframe displacements exceed its predictive range. At the same time, spatial overlap between consecutive frames approaches zero, rendering Intersection over Union (IoU)-based matching uninformative. Under these conditions, tracking becomes primarily appearance-driven. This work examines MOT behavior under extreme temporal sparsity and proposes Low-Frame-Rate Multi-Object Tracking (LFR-MOT), a purely appearance-based tracker that removes motion prediction entirely and relies on re-identification (ReID)-based appearance features with a two-stage matching strategy to handle detection uncertainty. Experimental results at 1 fps on UA-DETRAC and VisDrone show substantial improvements over motion-based and hybrid baselines. On VisDrone, existing motion-based methods yield Identity F1 (IDF1) as low as 17.3%, whereas LFR-MOT achieves 75.0%. On UA-DETRAC, LFR-MOT reaches 63.6% IDF1 under the same sparse temporal sampling. A separate within-dataset analysis on CityFlowV2 examines sensitivity to frame-rate reduction, and additional evaluations on real-world surveillance footage show that the method preserves identity consistency in challenging low-quality scenarios. Feature-based matching therefore provides a practical solution for surveillance systems operating under severe resource constraints and long interframe intervals.
Multi-object tracking (MOT) is an essential computer vision task that simultaneously tracks multiple objects in video sequences, with various applications in surveillance, autonomous navigation, and human-computer interaction. The tracking-by-detection (TBD) paradigm, which combines object detection with temporal assoc...
Yu-Jin Yang, Kyujin Shim, Kangwook Ko et al.· 0 citations
Real-time multi-object tracking systems remain highly vulnerable to full and long-term occlusion, where targets temporarily or completely disappear from the camera's field of view. Conventional trackers may terminate trajectories prematurely, resulting in identity loss and reduced situational awareness in applications...
Mais.M Mohammed, Sharifa Mohammed, Hanan Awadh et al.· 0 citations
The YOLOv8 family of models excels in object detection and demonstrates competitive performance relative to other methods; however, utilizing these models in resource-limited environments for real-time tracking presents challenges. This study introduces an innovative pedestrian tracking system that integrates a lightwe...
Sheeba Razzaq, Majid Iqbal Khan, Amil Roohani Dar et al.· IEEE Access· 0 citations
Diffusion-based detectors begin inference from noisy boxes, whereas tracking-by-detection pipelines usually
localize each video frame independently. This study examines whether propagated tracker boxes can guide a frozen
DiffusionDet detector without retraining. Confirmed tracks are propagated using their latest observ...
Muhammad Mustapha Miko, De-Quan Li· International Journal of Inn...· 0 citations
Multi-object tracking (MOT) has advanced rapidly in urban surveillance and autonomous driving, yet many trackers rely on ReID- and transformer-based appearance encoders and are designed for standard FoV cameras. These assumptions break down for low-cost omnidirectional deployments, where equirectangular projection intr...
Xin Shu, Meegan Gower, Y. Buckley et al.· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.