Skip to content
Conference

BigGaitMamba: Foundation-Driven State-Space Temporal Learning for Multi-View Gait Representation and Non-Cooperative Person Identification

Aug 2026 · International Conference Computational Vision and Bio Inspired Computing · pp. 923-928 · 0 citations · 17 references

Abstract

Gait biometrics support non-cooperative identification from distant surveillance video, yet viewpoint changes, clothing, carried objects, low resolution, occlusion, and long temporal dependencies reduce recognition reliability. BigGaitMamba addresses these conditions through a unified architecture combining foundation gait features, spatiotemporal attention, selective state-space temporal modeling, and masked video reconstruction. Projection adapters align independently pretrained modules within a common token space, while identity, metric-learning, and reconstruction objectives optimize a 256-dimensional descriptor. Evaluation on CASIA-B, OU-MVLP, Gait3D, and GREW produces 98.2% Rank-1 and 96.8% mAP, with precision, recall, and F1-score above $\text{9 6 \%}$. Component analysis attributes complementary gains to foundation initialization, attention refinement, long-range temporal modeling, and reconstruction learning. The resulting pipeline supports scalable multi-view gait retrieval across controlled and in-the-wild conditions.

View source

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.