A key strength of Proximal Policy Optimization (PPO) is its learned critic, which uses historical trajectories collected during reinforcement learning to estimate expected returns and reduce policy-gradient variance. However, we find that the critic is also a major source of instability in reinforcement learning for la...
Xuan-Yi Zhou, Qiu-Yang Mang, Huan-Zhi Mao et al.· 0 citations
Large language model (LLM) agents allocate test-time compute adaptively as they revise solutions, use tools, explore alternatives, and decide when to stop. This test-time strategy makes it difficult to measure how agent performance scales. We study open-ended tasks that provide continuous scores for intermediate submis...
Kai-Yuan Liu, Qiu-Yang Mang, Bo-Fei Peng et al.· 2 citations
MatrAIx is introduced, a population-scale simulated-user evaluation infrastructure for testing AI systems and digital products with heterogeneous users and provides an end-to-end infrastructure for evaluating AI systems and digital products with diverse simulated human users.
Xiaomin Li, Yuexing Hao, Jian Hou et al.· 1 citation
This work proposes TTCD, a TTT framework that introduces a self-supervised objective for allocating limited memory capacity for future use, and focuses on an in-place variant: In-Place TTCD, which uses the existing MLP parameters as the fast weights.
Zi-Xuan Wang, Xingyu Dang, Ruiming Zhu et al.· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.