HOMA-ST is presented, which recovers a complementary motion signal by decomposing the short-term dynamics of the bounding box into three physically meaningful branches and extracting a 21-dimensional high-order motion descriptor in which every dimension has an explicit interpretation and a known minimum-frame requirement.
Abstract
Single-object tracking from unmanned aerial vehicles (UAVs) is complicated by small target size, rapid camera ego-motion, and frequent occlusion, all of which degrade the appearance cues that Transformer trackers rely on. We present HOMA-ST, which recovers a complementary motion signal by decomposing the short-term dynamics of the bounding box into three physically meaningful branches — translation, geometric variation, and trajectory — and extracting a 21-dimensional high-order motion descriptor in which every dimension has an explicit interpretation and a known minimum-frame requirement. The three branches are fused by a branch-token self-attention block, and the resulting motion energy is injected as a learnable, gated residual on the search-side queries at every cross-attention layer. An auxiliary motion-consistency loss further couples spatial accuracy to temporal plausibility. On UAV123, UAVDT, DTB70, and VisDrone-SOT, HOMA-ST improves success AUC over OSTrack by 1.7 to 2.1%, with the largest gains on fast-motion and occlusion subsets. A further evaluation on 20 live-flight sequences recorded on a DJI Tello EDU platform confirms a 4.5% AUC improvement under real deployment conditions. The results indicate that HOMA-ST provides an effective framework for robust UAV tracking in both benchmark and real-flight scenarios.
Joint detection-and-embedding (JDE) trackers avoid per-detection crop inference by reading identity features for re-identification (ReID) from the detector. The detection-center readout, however, does not use a track prediction when forming the appearance descriptor. This is restrictive in unmanned aerial vehicle (UAV)...
Multiple object tracking (MOT) from unmanned aerial vehicles (UAVs) presents significant challenges due to drastic viewpoint changes and ambiguous top-down appearances. The former often result in large, misleading displacements of objects in the image plane, while the latter makes it difficult to distinguish between ob...
Ya-Xuan Hu, Jie Hua, Jian-Feng Ding et al.· IEEE Transactions on Geoscie...· 0 citations
Multi-object tracking (MOT) in UAV videos is a challenging task because small targets, partial occlusion, bounding-box displacement, and fluctuating detection confidence can reduce the reliability of frame-to-frame data association. In particular, potentially correct track-detection pairs may be rejected when a strict...
Heng-Yi Huang, Guang-Yao Zhou, Cheng-Han Yin et al.· 2026 3rd International Confe...· 0 citations
Airborne optical tracking of uncrewed aerial vehicle (UAV) swarms is challenging due to extremely small target scales, rapid viewpoint changes, and cluttered backgrounds, which can weaken target feature responses and lead to intermittent or temporarily missing detector responses. Existing multi-object tracking methods...
Zhao-Chen Chu, Tao Song, Ren Jin et al.· 0 citations
Aerial-ground cooperation requires real-time UAV--UGV relative-state information. Instead of maintaining global estimates for both robots, direct control in a UGV-attached non-inertial frame avoids reliance on global localization. Vision-based relative pose estimation with a passive marker offers a low-cost and effecti...
Ming-Xuan Zhang, Jia-Jun Yu, Bao-Zhe Zhang et al.· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.