Test-time Reinforcement Learning in Imperfect Information Games
This work extends the concept of gadget game, tabular technique for test-time search, to the reinforcement learning setting and formally proves that, unlike prior tabular algorithms, regularized policy-gradient algorithms limit possible strategy degradation caused by test-time reasoning, even without the gadget games.
Ondrej Kubícek, Viliam Lisý, Tuomas Sandholm
· 0 citations