Skip to content

Robots Acquire Manipulation Skills in Seconds from a Single Human Video

Jul 2026 · arXiv.org · Vol abs/2607.20033 · 2 citations · 107 references
Computer Science

TL;DR

HOST (Human-to-robot One-Shot Skill AcquisiTion), a framework that enables a robot to acquire skills in seconds from a single human video while retaining previously mastered skills.

Abstract

The ability to acquire skills rapidly and effortlessly while retaining those already mastered is essential for robots. However, current methods still rely on a cumbersome training-time loop that is costly and slow, while eroding skills already mastered. In this paper, we introduce HOST (Human-to-robot One-Shot Skill AcquisiTion), a framework that enables a robot to acquire skills in seconds from a single human video while retaining previously mastered skills. HOST resolves skill acquisition through a cascade of self-grounded prediction. It first estimates the robot's progress within the demonstrated task, then translates the upcoming progression into the robot's own future observations, and finally derives actions from these predicted observations. This cascade is trained on targets coupled to the video demonstration, obtained by mapping the robot trajectory and the video demonstration onto a shared task progress manifold, then redefining each target to align with the future progression of the video. HOST thereby enables the robot to actively follow the demonstrated procedure and adapt it to the robot's embodiment. HOST acquires novel skills at inference time from a single human video in an average of 29 seconds and achieves a 62% average success rate. It exceeds the zero-shot baseline by 45% while retaining previously mastered skills. HOST even exceeds the baseline fine-tuned on 50 robot demonstrations per task while requiring 50 times fewer demonstrations and acquiring each skill 507 times faster. Additional information about HOST is available on the project website.

View source

Similar papers

Preprint Aug 2026

Robot Learning from Human Demonstrations: Handwritten Alphabet Trajectories and Human-Likeness Evaluation

A framework for learning human-like robot motion from demonstration, including data collection, probabilistic trajectory learning, and perceptual user evaluation is presented, extending the widely used Gaussian Mixture Model and Gaussian Mixture Regression approach for learning from demonstration.

Alperen Kenan, Paul A. Bremner, Manuel Giuliani · 1 citation
Jul 2026

Learning Forward & Reverse Skills from a Single Unfinished Demonstration for Constrained Manipulation Tasks

Learning from demonstration (LfD) enables robots to learn manipulation skills directly from expert demonstrations but remains challenging for contact-rich tasks involving geometric constraints and force interaction. Existing approaches typically require multiple complete demonstrations and do not support reverse skill execution. In this paper, we present a unified one-shot framework for constrained manipulation that learns both forward and reverse execution from a single, possibly unfinished demonstration. Our method decomposes demonstrations into non-contact and contact phases, with non-contact motion encoded with dynamic movement primitives (DMP), and contact motion represented as a sequence of screw motion primitives segmented by our proposed geometry-driven twist-direction segmentation algorithm. During execution, screw primitives are executed sequentially under admittance-guided pose correction and speed regulation, enabling task completion beyond the demonstrated trajectory length as well as reverse skill execution without additional learning data. Experiments on peg insertion, battery insertion, lock opening, and screw driving tasks demonstrate improved success rates and robustness over segmentation and one-shot trajectory learning baselines. Details are available on the project website: https://tuwien-asl.github.io/LfD-Screw/.

Yexin Hu, Haoyi Zheng, Johannes Heidersberger et al. · 0 citations
Open access

Active Robot Learning from Demonstrations with Human Teachers

This thesis investigates Active Learning from Demonstrations (Active LfD) as a principled approach to reduce distribution shift while accounting for human factors in realistic human-in-the-loop settings and develops an Active LfD algorithm that operates with offline demonstrations and guides the resulting demonstration distribution toward a more balanced one, or more generally toward any specified target distribution, so as to improve robot learning.

Muhan Hou · 0 citations
Review Jul 2026

Tri-Manual Visuomotor Imitation Learning of Robot Policies

TriManPolicy is presented, a tri-manual imitation learning system that allows one operator to demonstrate behaviours for three arms while reconsidering when they occur, and policies trained on demonstrations retimed by DATS exhibit more efficient coordination while maintaining comparable observed task success.

James Zhao, Mingyuan Ba, Weiming Zhi · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.