Terminal-using agents benefit from reinforcement learning (RL) in coding, debugging, and other multi-step terminal tasks. In these tasks, later commands often depend on information or intermediate results produced by earlier commands. However, existing trajectory-level and step-level credit assignment methods do not ex...
Yu Li, Guang-Feng Cai, Long-Fei Li et al.· 0 citations
ANIMASK, a simulation framework that freezes books and scripts into story worlds whose characters act on their own motivations and replays each story from its freeze point, is introduced.
Xiu-Cheng Zhang, Zhuo-Ning Xu, Han-Jun Luo et al.· 0 citations
MasDrift exposes a centralization tradeoff and makes authorization preservation a measurable property of MAS design, and a heterogeneous case study confirms that the failure follows from coordination rather than model strength.
Zhuo-Ning Xu, Xiu-Cheng Zhang, Han-Jun Luo et al.· 2 citations
This work instantiates AutoCRAT, a decoder-side controller for frozen backbones that operates over a discrete action space and updates control decisions only at semantic boundaries, improving stability while remaining responsive to the evolving reasoning process.
Han-Jun Luo, Qiu-Shi Liu, Jing-Yang Zhang et al.· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.