In security-sensitive applications, language-model agents are often required to coordinate without disclosing confidential information. Yet repeated interactions may also let ordinary messages acquire shared private meaning. We study a repeated game with pairs of models in which the sender model observes one of four se...
Jacob Dineen, Si-Lei Ren, Mu-Hao Chen et al.· 0 citations
BOW is introduced, an RL framework that instead trains models to produce self-contained, neutral, and comprehensive descriptions of the plausible next-word space, and human evaluation shows that BOW-Reg produces broader next-word reasoning trajectories, while direct next-word-prediction evaluation shows that these traj...
This work proves a PAC-Bayes bound guaranteeing that a dictionary extracted from successful trajectories has bounded expected description length on future successful behavior, and introduces ReuseRL, which grounds agentic RL in the Minimum Description Length (MDL) principle.
Zhikun Xu, Yu Feng, Jacob Dineen et al.· arXiv.org· 1 citation
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.