Skip to content
Open access

Anomaly Detection and Recognition in Complex Commercial Kitchen Environments

2026 · IEEE Access · Vol 14, pp. 117929-117945 · 0 citations · 30 references

TL;DR

A complete framework that combines data construction and detection model enhancement, developed by integrating three complementary modules from prior studies, indicates that the proposed method provides a practical solution for complex kitchen anomaly detection and intelligent food-safety monitoring.

Abstract

Commercial kitchen surveillance provides important visual evidence for food-safety supervision, but automatic anomaly detection in such scenes is still challenging. Abnormal events are usually sparse and are embedded in cluttered operating environments with occlusion, low illumination, and large appearance variation. This study focuses on three representative anomalies: rat intrusion, staff smoking, and staff upper-body clothing violation. These categories cover two different recognition difficulties. Rat intrusion requires reliable tiny-object detection, whereas the two staff-related categories require the model to capture subtle local cues while also using human posture and surrounding scene context. To address these issues, this paper develops a complete framework that combines data construction and detection model enhancement. For data construction, a semi-automatic annotation-assistance workflow is built using SAM3 and Qwen3-VL-32B. SAM3 recalls candidate regions. Qwen3-VL-32B verifies candidate categories using local patches, global context, and task-specific prompts. Before manual verification, the SAM3–Qwen3-VL workflow achieves 0.928 candidate recall and 0.903 label precision. Human verification further improves the final candidate recall and label precision to 0.971 and 0.976, respectively. The complete annotation workflow requires only 38.2% of the time used by fully manual annotation. For detection, a task-oriented YOLOv11 adaptation is developed by integrating three complementary modules from prior studies. FeaturePyramidSharedConv is used to enhance high-level multi-scale context, MultiScaleGatedAttn is used to strengthen adaptive cross-layer feature selection, and DynamicScalSeq is used to reinforce the $P_{3}$ small-object branch through stacking along the scale dimension and max-response selection. Experiments on a custom kitchen anomaly dataset show that, under the default YOLOv11n setting with an input size of $640\times 640$ , the improved model achieves 0.862 mAP@0.5, which is 3.3 percentage points higher than the YOLOv11n baseline. Additional experiments further show that each adapted component contributes to the final performance, and that the framework remains effective under different MSGA placements, DynamicScalSeq variants, input resolutions, and model scales. These results indicate that the proposed method provides a practical solution for complex kitchen anomaly detection and intelligent food-safety monitoring.

Read PDF

Similar papers

Open access Jul 2026

Enhanced deep learning model for anomaly object detection and tracking from surveillance videos.

Video anomaly detection plays a crucial role in video surveillance, which identifies suspicious intruders without human intervention. Moreover, the rapid growth of video surveillance applications such as intrusion detection, health monitoring systems, and fault detection provides a secure environment. Furthermore, detecting anomalous intruders from video is a challenging task because of diverse contexts, lack of training data, and environmental variations. Several conventional techniques use various Deep Learning algorithms for anomaly detection, which possess limitations including high false positive rates and occlusion. Therefore, to overcome the drawbacks, efficient anomaly object detection and tracking system is proposed using an enhanced wolf Crocuta optimization-based deep Bidirectional Long Short-Term Memory (EnWC-DBiLSTM) classifier. Here, an effective keyframe selection is attained by the Timber Prairie Wolf Optimization (TPWO) strategy, which optimally selects the required keyframes for further processing. Further, the combination DBiLSTM classifier processes the input data accurately and detects the target object, in both directions. Moreover, the enhanced Wolf Crocuta optimization (EnWC) helps to eliminate local power resolution, which improves the convergence speed of the model. Henceforth, the proposed model achieved an accuracy of 98.226%, equal error rate, sensitivity, and specificity of 1.774, 98.111%, and 99.551%, respectively, for the ShanghaiTech campus dataset.

B. Gayal, S. Patil, D. Meshram et al. · 0 citations
Preprint Jul 2026

Context-structured Video Anomaly Detection with Large Vision-Language Models

Training video anomaly detectors is challenging due to the difficulty and cost of annotating diverse and rare abnormal events. Although recent large vision-language models enable training-free inference, existing approaches mostly rely on holistic inference over sampled video and may miss context-specific anomaly cues. In this paper, we present CSI-VAD, a training-free video anomaly detector that identifies abnormal events across diverse contexts. The key idea is to decompose each video into three distinct contexts (environment, objects, time) and perform context-specific inference in separate branches. Because we ground anomaly judgments solely in context-specific visual cues, we do not require predefined text prompts describing abnormal events or dataset-specific tuning. Experiments on UCF-Crime and UBnormal show that CSI-VAD consistently improves over the direct holistic baseline and achieves competitive performance against existing methods, showing the advantage of structured context decomposition for training-free video anomaly detection.

Dongjun Kim, Changjae Oh, Andrea Cavallaro et al. · 0 citations
Open access Aug 2026

Object-Centric Industrial Video Anomaly Detection with Local-Global Representation Learning

Industrial video anomaly detection is a critical component of modern smart manufacturing and industrial surveillance, aiming to automatically identify deviations from normal operational patterns. Traditional methods relying on frame-level or pixel-level feature extraction often struggle with the complex, dynamic, and heavily occluded environments typical of industrial settings. This paper proposes a novel framework centered on object-centric video anomaly detection, heavily augmented by a local-global representation learning mechanism. By shifting the analytical focus from the entire image frame to specific objects of interest, such as machinery components, manufactured goods, and human operators, the proposed method isolates highly relevant features while mitigating the impact of background noise and illumination variations. The framework utilizes a robust tracking-by-detection paradigm to construct spatio-temporal object tubes, which are subsequently processed to extract localized representations capturing appearance and motion dynamics. Concurrently, a global representation module employs attention mechanisms to model the complex interactions between multiple objects and their contextual environment. The integration of these local and global streams ensures a comprehensive understanding of the industrial scene, allowing for the precise localization and classification of anomalous events. Extensive evaluations on multiple large-scale industrial datasets demonstrate that the proposed object-centric framework significantly outperforms existing state-of-the-art approaches in both anomaly detection accuracy and computational efficiency. The findings suggest that integrating structured object interactions into representation learning provides a highly scalable and robust solution for real-world industrial monitoring. 

Chun-Wai So, Man-Kit Chau · 0 citations
Open access Aug 2026

GMS-YOLO11n: A Sheep Detection Model for Challenging Fixed-View Farm Conditions Integrating Spatially Gated Structural Enhancement and Multi-Scale Attention

Simple Summary Sheep detection is an important part of intelligent farm management because it supports automatic counting, individual tracking, behavior observation, and health monitoring. However, accurately detecting sheep can be difficult when animals gather closely together, partially block one another, appear at different distances from the camera, or are recorded under poor lighting and complex backgrounds. To address these challenges, we developed GMS-YOLO11n, a new computer vision model for automatic sheep detection. Our approach improves the original YOLO11n model by strengthening important visual features, such as sheep edges, body contours, and local textures, while combining information from different image scales. In the current within-farm evaluation, the proposed model reduced missed detections and improved bounding-box localization, particularly in the nighttime low-light and obvious-occlusion subsets. These findings are limited to the fixed-view surveillance conditions represented in the present dataset and should not be interpreted as evidence of general performance across different farms, breeds, camera devices, or viewpoints. External validation is required before the model can be considered suitable for broader farm deployment.

Wenbo Yu, Ruoya Xie, Yongqi Liu et al. · 0 citations
Open access Jul 2026

IntelligentVehicle Security: Real-Time Anomaly Detection and Anti-Theft Surveillance Using Monocular Depth Estimation and Behavioral Analysis

A proactive, real-time computer vision system designed to detect potentially suspicious behavior around parked vehicles, with a specific focus on unauthorized proximity and loitering is proposed, making it a strong candidate for practical urban vehicle monitoring, subject to further large-scale validation across diverse environments.

Umar Adeel, Ammar Rashid, S. Yusof et al. · 0 citations
Aug 2026

A Multi-Scale Attention-Enhanced Algorithm for Hand Keypoint Detection

Hand keypoint detection is important for human – computer interaction and industrial process monitoring, but practical deployment on resource-constrained devices still faces challenges such as the trade-off between accuracy and efficiency, limited robustness in dynamic scenes, and sensitivity to occlusion. To address these issues, this paper presents a two-stage hand analysis framework that combines an optimized YOLOv5s detector with HRNet-based keypoint estimation. In the detection stage, the backbone is replaced with InceptionNeXt and an adapted MSA-CAM module is introduced to improve feature representation in cluttered industrial scenes while reducing computational cost. In the pose stage, HRNet is used to estimate 21 hand keypoints from detected hand regions. Experiments on multiple hand datasets show that the proposed detector achieves a favorable balance between accuracy and efficiency. In a discrete workshop packaging scenario, the overall system also supports action-sequence recognition and anomaly detection, achieving 96.3% recognition success in the topview setting and 95.0% in the front-view setting. These results demonstrate the practical value of the proposed framework for real-time industrial hand analysis.

Longxin Lin, Xin Wang, Deping Li · 0 citations