Skip to content

Author

Guangqian Kong

1 paper indexed here

We haven’t gathered this author’s papers yet. Follow them and we’ll fetch their work.

Not the right person? Other researchers publish under this name.

Jul 2026

Spatiotemporal transformer with cross-temporal gated fusion for single object tracking

Existing single-object trackers often struggle to fully exploit temporal information and distinct feature representations, limiting their robustness in complex scenarios. Specifically, current approaches face challenges in (1) capturing channel-wise discriminative features during target deformation, (2) adapting to sudden appearance changes due to reliance on static memory banks, and (3) mitigating background interference within transformer attention mechanisms. To address these issues, we propose STFTrack, a proposed tracking framework integrating enhanced spatiotemporal features. First, we construct a semantic–spatial–instance attention module, which refines target representation via cascaded channel calibration and instance uncertainty modeling. Second, a gated attention fusion module is introduced to adaptively aggregate temporal and spatial cues. Third, we design an improved spatiotemporal decoder equipped with learnable autoregressive queries to establish a robust cross-frame state propagation mechanism. Extensive experiments on five benchmarks, including LaSOT, GOT-10k, and TrackingNet, demonstrate the superiority of STFTrack. For instance, it achieves an area under the curve score of 72.8% on the LaSOT benchmark. Qualitative results further confirm that STFTrack effectively suppresses background clutter and maintains accurate tracking under challenging conditions including deformation, fast motion, and occlusion.

Xun Duan, Guangqian Kong · 0 citations