This work aims to establish a unified perspective on vision-based mistake analysis in procedural activities, highlighting its potential across diverse domains and aspects and categorizing approaches based on their use of procedural structure, supervision levels and learning strategies.
Action recognition has emerged as a critical area of research within the realm of computer vision, driven by the increasing demand for intelligent human–machine systems capable of understanding and interpreting human behaviors in the real world. The ability to decipher intricate details of human actions holds immense potential to improve system design, predictive modeling, data-informed decision-making, and real-time operational improvements across a wide variety of domains. Some examples of applications range from surveillance and real-time management of public spaces and infrastructure systems, to development of predictive modeling and robotic systems for individualized healthcare interventions, to implementing effective human–computer interaction in both professional and recreational settings. This paper provides a comprehensive survey of the current state of action recognition, focusing specifically on three open-world challenges: the integration of multimodalities, the ethical and social implications of these technologies, and the utilization of feedback mechanisms to enhance model performance. We delve into the evolution of action recognition, from early feature-based approaches to the deep learning revolution, emphasizing how the incorporation of multiple sensory modalities—such as visual, audio, and depth data as well as other cues—has advanced the field. Furthermore, we examine the ethical challenges associated with deploying these technologies in the public domain, particularly regarding privacy, bias, and societal impact, and discuss the need for responsible development and regulation. The third focus of the paper is the use of top-down and bottom-up feedback mechanisms within deep learning architectures, exploring how these strategies can mimic human cognitive processes to improve accuracy and reliability in action recognition systems. By identifying current gaps and proposing future research directions, this paper aims to inspire continued innovation in this dynamic and impactful field for intelligent systems.
An overview of Vision LLM architectures, their applications and the challenges they face and case study of how building of AI Models through visionLLM may help IndoAI AI camera system are provided.
Rohit Yadav· Journal of Artificial Intell...· 0 citations
Errors are inevitable in procedural tasks, yet most AR guidance systems focus on step-by-step instruction delivery rather than helping users recognize and recover from mistakes. We present RegulAR, an AR task assistant for procedural error recognition and recovery. RegulAR models task instructions as a hierarchical dependency graph and combines this structure with a Multimodal Large Language Model (MLLM) to interpret egocentric observations during execution. This enables RegulAR to track progress, identify deviations by error type, estimate their impact on later steps, and deliver appropriately salient interventions through an in-situ head-up display that visualizes task state and recovery guidance. By making procedural structure explicit, RegulAR supports not only next-step guidance, but also reasoning about what went wrong, why it matters, and how users can get back on track. In a within-subject study (N=12), participants reported better task-structure understanding and recovery support with RegulAR than the MLLM-only baseline.
Yi-Lin Ye, Jin-Du Wang, Hiu-Tung Wong et al.· 0 citations
IMBENCH is introduced, a benchmark designed to evaluate intuitive manipulation as an integrated capability spanning perception, physical reasoning, action generation, and iterative execution, and position IMBENCH as a step toward evaluating and enabling more integrated, adaptive physical intelligence.
Anurag Maurya, Sukhvansh Jain, Prajwal Avhad et al.· 0 citations
Visual perception plays a critical role in industrial assembly systems, where robotic actions are driven by image-based sensing under variable and data-dependent conditions. A key challenge in such systems lies in transforming unstructured visual inference outputs, including object detection and pose estimation results, into structured and executable control parameters that can be reliably grounded in physical execution. In practice, mismatches between perception outputs and downstream control interfaces often lead to execution errors and reduced system robustness in multi-stage assembly processes. To address this challenge, this paper proposes a vision-guided robotic action generation framework that explicitly models the data flow from visual perception to robotic action execution. The framework introduces a structured visual data extraction mechanism that interprets raw visual outputs into type-consistent, constraint-aware, and physically feasible motion parameters, enabling reliable perception-action coupling in industrial assembly systems. By decoupling visual interpretation from low-level control execution, the proposed approach improves modularity and robustness across heterogeneous hardware platforms and low-code industrial orchestration environments. The proposed framework is implemented and evaluated through an end-to-end, data-dependent toy vehicle assembly task involving multiple perception-driven operations. Experimental results demonstrate that the proposed method significantly improves perception-action alignment robustness, achieving higher phase-level execution reliability and an end-to-end assembly success rate of up to 94%, outperforming baseline approaches that lack explicit visual data alignment mechanisms.
Longxiang Huang, Jiaxin Dai, Tao Wang et al.· International Conference on...· 0 citations
ALVA (Action- and Language-Conditioned Video Assessment), a trajectory evaluator that conditions its assessment on visual observations, the executed action sequence, and the natural language instruction, provides more effective feedback than the evaluated static image and embedding-based visual baselines and reduces the performance gap to a ground-truth oracle.
Hwanhee Kim, Jaehyun Jang, Seung-Min Cha et al.· Italian National Conference...· 0 citations