Skip to content
#small language model Conference Open access

A Review of Multimodal Perception and Closed-Loop Task Execution for Human-Robot Collaboration Robots Oriented Toward Intelligent Assembly

Sep 2026 · Exploring Science Academic Conference Series · 0 citations · 31 references

TL;DR

This review discusses recent research from the perspective of a complete "perception-decision-execution-feedback" process, and focuses on the main topics operator action and intention recognition, perception of part states, task allocation under human-related constraints, safe trajectory planning, and quality feedback.

Abstract

With the increasing demand for flexible manufacturing, especially for multi-variety and small-batch production, human-robot collaboration has become more widely used in intelligent assembly. Many studies have investigated multimodal perception, task alloca tion, safety control, and quality inspection, and considerable progress has been made in these areas. Even so, several problems are still difficult to address. Information exchange between different technical modules is often not well connected, task under standing is still limited, and feedback from the assembly process is not fully incorporated into decision making. This review discusses recent research from the perspective of a complete "perception-decision-execution-feedback" process. The main topics inc lude operator action and intention recognition, perception of part states, task allocation under human-related constraints, safe trajectory planning, and quality feedback. Rather than considering these techniques separately, attention is given to how they interact during collaborative assembly. Current studies suggest that the main challenge is no longer improving the accuracy of a single perception or control module. A more important issue is how to connect multi-source perception, task understanding, safety constraints, and quality feedback into a stable closed-loop framework. Future work is expected to pay more attention to cross-station adaptation with limited training samples, practical and verifiable use of vision - language-action models, digital twin-based safety validation, and skill updating based on assembly deviations. It is hoped that this review can provide useful ideas for future research and system development in intelligent human-robot collaborative assembly.

Read PDF

Similar papers

#reinforcement learning Review Open access Sep 2026

Review of key technologies for robot embodied intelligence oriented toward flexible manufacturing

The review finds that multimodal fusion and semantic SLAM are overcoming perception bottlenecks, that deep learning and force/position hybrid control are balancing flexible adaptability with high-precision operation, and that deep reinforcement learning and large models are advancing intelligent process planning.

Zheng-Yang Chen · 0 citations
Open access Sep 2026

Action Recognition in Human-robot Teaming for Assembly Tasks

This paper presents a study on Human-Robot Teaming (HRT), focusing on state-of-the-art methodologies and real-world implementation of a collaborative system for industrial assembly tasks. The study explores the integration of real-time human action recognition using deep learning, specifically LSTM (Long Short-Term M...

Federico Neri, Alessandro Frezza, Giacomo Palmieri et al. · 0 citations
Preprint Sep 2026

RoboAssist: Interactive Human-Humanoid Planning for Long-Horizon Surgical Assistance

Long-horizon surgical assistance requires humanoid robots to coordinate with evolving human activities while maintaining safety across planning and execution. We present RoboAssist, an agent-based framework for interactive human-humanoid planning that integrates workflow reasoning, task coordination, and cross-layer sa...

Jing-Wei Jia, Ke-Yu Zhou, Jie-Wei Wang et al. · 0 citations
Aug 2026

MulPlanLM: multimodal robotic task planning with vision-language models and physical feedback

Experimental results in various task scenarios show that the proposed framework consistently improves overall task success rates compared with unimodal settings with different LLMs and achieves a higher success rate compared to using only visual or force data.

Young-Chae Son, Dong-Han Lee, Soo-Chul Lim · 0 citations

Related blog posts

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.