Goal-conditioned reinforcement learning (GCRL) relies heavily on how target goals are represented to the policy. While recent methods encode goals via temporal distance, occupancy, or controllability, it remains unclear how much downstream performance actually depends on representation quality. We study this in offline...
Syed Nazmus Sakib, Abdul Monaf Chowdhury, Nafiul Haque et al.· 0 citations
Reinforcement learning with verifiable rewards (RLVR) has become an important approach for improving reasoning during post-training. Recent work suggests that some difficult prompts remain resistant to learning even when they occasionally produce correct solutions. We revisit this unlearnability phenomenon and find tha...
Grounded Transitive RL (GTRL), an offline GCRL value learning algorithm that grounds the divide-and-conquer update with a one-step TD target, and corrects the bias from hindsight relabeling by reweighting each goal against how reachable it was from other successors.
Abdul Monaf Chowdhury, Sameer Iqbal Chowdhury, Shifat E. Arman et al.· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.