On-policy distillation (OPD) supervises a student on its own trajectories with token-level signals from a frozen teacher, yet how a sampled loss allocates updates across tokens remains poorly understood. We analyze the gradient of the per-token K2 estimator of reverse KL with respect to the student logits. The $\ell_1$...
Bing Shao, Jia-Zheng Zhang, Long Ma et al.· 0 citations
As AI agents become participants in the development of their successors, they reshape both the production of intelligence and the role of human researchers. We introduce Atria Dawn Preview, a foundation agentic language model designed for scientific research and engineering workflows, with the goal of expanding the fro...
The Prefix-Adaptive Block Diffusion Model (PA-BDM) is proposed, which replaces intra-block bidirectional denoising with causal denoising from prefix to suffix and treats the block size as a maximum candidate range rather than a fixed commitment unit.
Ming-Xu Chai, Zi-Yu Shen, Chen-Yu Liu et al.· arXiv.org· 0 citations
A novel textual representation of fault trees is proposed, and a benchmark for multi-turn dialogue systems that emphasizes robust interaction in complex environments is constructed, evaluating a model's ability to assist in malfunction localization.
Yuhui Wang, Zhi-Xiong Yang, Ming Zhang et al.· arXiv.org· 0 citations
CAFE (Coupled Agent--Feedback Evolution), a framework in which a shared-parameter model alternates between search-agent and critic roles, is introduced, suggesting that a self-improving search agent needs feedback that co-evolves with the policy it guides.
Bo-Yang Liu, Senjie Jin, Pei-Xin Wang et al.· 1 citation
To mitigate a critical imbalance during the exploration-and-learning process, this work approaches head-tail re-balance during the exploration-and-learning process from two perspectives: distribution-reshaping and trajectory-resampling.
Xin Guo, Zhiheng Xi, Yiwen Ding et al.· Annual Meeting of the Associ...· 1 citation
AgentGym2 is presented, a new evaluation framework with task instances grounded in real-world end-to-end working demands that measures agents'ability to execute end-to-end procedures, discover tools via exploration, compose tools for unseen tasks, and remain robust to noisy and underspecified information.
Zhiheng Xi, Dingwen Yang, Jiaqi Liu et al.· Annual Meeting of the Associ...· 1 citation
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.