Skip to content

Generative Adversarial Self-Imitation Learning With Large Language Model Feedback for Robot Control and Navigation

2026 · IEEE Transactions on robotics · Vol 42, pp. 2877-2897 · 0 citations · 58 references
Computer Science

Abstract

Deep reinforcement learning (DRL) has achieved great success in many simulated and real-world robotic tasks. However, the difficulty of designing efficient and dense reward functions makes applying DRL to tackle complex long-horizon and open-world tasks a great challenge. Generative adversarial imitation learning (GAIL) can directly learn policies from the expert trajectories and generalize well in large and complex environments, but relies on high-quality demonstrations and can seldom surpass the performance of the demonstration. Recent work used additional human evaluative feedback to facilitate GAIL to learn faster and surpass the demonstrations, but still requires suboptimal demonstrations. Moreover, it is costly and difficult for human expert to provide relatively high-quality demonstrations and evaluative feedback for various tasks. To address the above issues, in this article, we propose generative adversarial self-imitation learning from demonstration and large language model (LLM) feedback (GASL<inline-formula><tex-math notation="LaTeX">$^{3}$</tex-math></inline-formula>MF), since LLMs encode rich commonsense knowledge and can perform a variety of reasoning tasks. GASL<inline-formula><tex-math notation="LaTeX">$^{3}$</tex-math></inline-formula>MF allows a robot to learn from poor demonstrations and gradually replace them with its own good trajectories evaluated by LLM feedback. Our results in four physics-based control tasks and a mobile robot navigation task show that, even with demonstrations of poor performance or not completing the task, GASL<inline-formula><tex-math notation="LaTeX">$^{3}$</tex-math></inline-formula>MF can learn faster with close to optimal performance, and generalize well to different environments and the real world with sim-to-real adaptation. Further analysis shows that the overall distribution of LLM feedback closely resembles that of human feedback and remains closer to that of ground-truth rewards than human feedback. Finally, our GASL<inline-formula><tex-math notation="LaTeX">$^{3}$</tex-math></inline-formula>MF method works regardless of the LLM employed, and the LLM feedback from different LLMs remain robust across tasks and even better consistency than human feedback for robot learning in some tasks. These results shed light on the potential of robot imitation learning from even poor or failed demonstrations and broaden its application to a wide range of real-world tasks.

View source

Similar papers

Preprint Jul 2026

Discriminative Barrier Functions for Safe Adversarial Imitation Learning from Observation

This work constraining reward function candidacy during IRL to the space of CBFs yields a formulation that exhibits safe online control with continuous experiential improvement, and demonstrates that the recovered barrier function is robust to unsafe states entirely absent from the expert data.

Anubhav Vishwakarma, Bhaumik Mehta, Caleb Hsu et al. · 0 citations
Preprint Aug 2026

Efficient Real-World Online Reinforcement Learning for Robot Manipulation via Centralized Training and Critic Decomposition

A unified framework that combines centralized training with decentralized execution (CTDE) and a Hybrid Reward Architecture (HRA) is introduced that enables multiple actors to share a centralized multi-head critic and substantially improves both sample efficiency and policy performance.

Changhao Li, Yifang Zhang, Heng Zhang et al. · 0 citations
Preprint Jul 2026

Learning Reach-Avoid Task with Reinforcement Learning: Vectorized Simulation and Benchmark

This work identifies that additional research is still required to claim the successful resolution of the robotic arm reach-avoid task using DRL, and presents a comprehensive benchmark for the reachavoid task that accurately captures real-world complexities without simplifications.

Jonas Weihing, Shahram Eivazi · 0 citations
2025

Multi-Agent Imitation by Learning and Sampling from Factorized Soft Q-Function

Learning from multi-agent expert demonstrations, known as Multi-Agent Imitation Learning (MAIL), provides a promising approach to sequential decision-making. However, existing MAIL methods including Behavior Cloning (BC) and Adversarial Imitation Learning (AIL) face significant challenges: BC suffers from the compounding error issue, while the very nature of adversarial optimization makes AIL prone to instability. In this work, we propose M ulti-A gent imitation by learning and sampling from F actor I zed S oft Q-function (MAFIS), a novel method that addresses these limitations for both online and offline MAIL settings. Built upon the single-agent IQ-Learn framework, MAFIS introduces the value decomposition network to factorize the imitation objective at agent level, thus enabling scalable training for multi-agent systems. Moreover, we observe that the soft Q-function implicitly defines the optimal policy as an energy-based model, from which we can sample actions via stochastic gradient Langevin dynamics. This allows us to estimate the gradient of the factorized optimization objective for continuous control tasks, avoiding the adversarial optimization between the soft Q-function and the policy required by prior work. By doing so, we obtain a tractable and non-adversarial objective for both discrete and continuous multi-agent control. Experiments on common benchmarks including the discrete control tasks StarCraft Multi-Agent Challenge v2 (SMACv2), Gold Miner, and Multi Particle Environments (MPE), as well as the continuous control task Multi-Agent MuJoCo (MaMuJoCo), demonstrate that MAFIS achieves superior performance compared with baselines. Our code is available at https://github.com/LAMDA-RL/MAFIS .

Yichen Li, Zhongxiang Ling, Tao Jiang et al. · 3 citations
Preprint Jul 2026

Dynamics-Aware Meta-Imitation for Generalization to Unseen Robotic Manipulation

Imitation Learning aims to learn skills from extensive observations and demonstrations for robots, so it suffers from data scarcity and environment generalization. The existing methods predominantly focus on imitation from in-domain tasks and consequently struggle with generalization to unseen tasks. To bridge this generalization gap, we propose the \textbf{D}ynamics-\textbf{A}ware \textbf{M}eta-\textbf{I}mitation (DAMI) framework. By integrating meta-learning to construct a shared skill space, DAMI equips agents for rapid adaptation to novel tasks. We introduce the Visual-Motor Trajectory (VMT) module to capture complex spatio-temporal dynamics within the task latent space. Furthermore, we propose the Unpaired Unified Task (U2T) block to fuse unstructured multimodal observations. To coordinate these representations, we integrate a Task-Conditioned Feature Modulation (TCFM) mechanism customized for modulating low-level 3D features. By capturing intrinsic dynamics from a random complete reference demonstration, our framework learns the underlying task logic rather than memorizing static cues, ensuring effective generalization. Extensive experiments in both simulation and real-world settings demonstrate that our approach outperforms state-of-the-art baselines regarding direct inference on seen tasks and adaptation to unseen tasks via few-shot fine-tuning.

Zhenduo Shang, Xiyao Liu, Bohan Li et al. · 0 citations
Conference Jul 2026

Hybrid Large Language Model-Reinforcement Learning Pipeline to Enhance Simulation Training for Robotics

Despite rapid advances in artificial intelligence, robotic systems remain limited by poor generalisation across unstructured environments and fragile training pipelines. Reinforcement learning (RL) has shown promise in training robotics, yet its effectiveness is often constrained by manually engineered reward mechanisms. In parallel, large language models (LLMs) demonstrate strong reasoning and evaluation capabilities that remain underutilised in robotic training pipelines. This paper proposes a hybrid LLM-RL framework in which an LLM dynamically evaluates robot performance during simulation training and adaptively modifies the reward weights to improve learning stability, accuracy of task completion, and policy convergence. Unlike existing work that focuses on natural language control at inference time, the proposed method leverages the LLM during training, acting as a high-level reward critic. We implemented this framework using an open-source robotic arm trained in simulation to demonstrate improved task success rates and learning efficiency compared to static reward mechanisms. This work highlights a scalable pathway toward more adaptive and generalisable robotic training systems for advanced robotics.

Parith Avasadanond, Jovan Hartono, Kenneth Y. T. Lim · 0 citations