Temporal Expert-Based Reward Learning for Inverse Reinforcement Learning
Abstract
Learning reward functions from expert demonstrations removes the need for manual reward engineering in reinforcement learning applied to robotic manipulation. Existing methods, however, require trajectory quality annotations, episode success labels, or produce implicit rewards that are difficult to inspect. This paper presents TEXB-IRL (Temporal EXpert-Based reward learning for Inverse Reinforcement Learning), which learns a dense neural reward function directly from unlabeled expert demonstrations. TEXB-IRL exploits intra-trajectory temporal ordering through two complementary objectives: a temporal consistency loss that enforces monotonically increasing reward along expert trajectories, and an expert-agent separation loss that anchors the reward scale. The policy is optimized with PPO over the learned reward. Preliminary experiments in simulation with the PAL TIAGo++ robot show competitive results against state of the art algorithms.