Skip to content

Author

Ondrej Kubícek

1 paper indexed here

We haven’t gathered this author’s papers yet. Follow them and we’ll fetch their work.

Not the right person? Other researchers publish under this name.

#small language model Preprint Aug 2026

Test-time Reinforcement Learning in Imperfect Information Games

This work extends the concept of gadget game, tabular technique for test-time search, to the reinforcement learning setting and formally proves that, unlike prior tabular algorithms, regularized policy-gradient algorithms limit possible strategy degradation caused by test-time reasoning, even without the gadget games.

Ondrej Kubícek, Viliam Lisý, Tuomas Sandholm · 0 citations