Object-Centric Industrial Video Anomaly Detection with Local-Global Representation Learning
Industrial video anomaly detection is a critical component of modern smart manufacturing and industrial surveillance, aiming to automatically identify deviations from normal operational patterns. Traditional methods relying on frame-level or pixel-level feature extraction often struggle with the complex, dynamic, and heavily occluded environments typical of industrial settings. This paper proposes a novel framework centered on object-centric video anomaly detection, heavily augmented by a local-global representation learning mechanism. By shifting the analytical focus from the entire image frame to specific objects of interest, such as machinery components, manufactured goods, and human operators, the proposed method isolates highly relevant features while mitigating the impact of background noise and illumination variations. The framework utilizes a robust tracking-by-detection paradigm to construct spatio-temporal object tubes, which are subsequently processed to extract localized representations capturing appearance and motion dynamics. Concurrently, a global representation module employs attention mechanisms to model the complex interactions between multiple objects and their contextual environment. The integration of these local and global streams ensures a comprehensive understanding of the industrial scene, allowing for the precise localization and classification of anomalous events. Extensive evaluations on multiple large-scale industrial datasets demonstrate that the proposed object-centric framework significantly outperforms existing state-of-the-art approaches in both anomaly detection accuracy and computational efficiency. The findings suggest that integrating structured object interactions into representation learning provides a highly scalable and robust solution for real-world industrial monitoring.