GTRL: Grounding Divide-and-Conquer Value Learning with Temporal Differences
Grounded Transitive RL (GTRL), an offline GCRL value learning algorithm that grounds the divide-and-conquer update with a one-step TD target, and corrects the bias from hindsight relabeling by reweighting each goal against how reachable it was from other successors.
Abdul Monaf Chowdhury, Sameer Iqbal Chowdhury, Shifat E. Arman et al.
· 0 citations