On-policy distillation (OPD) supervises a student on its own trajectories with token-level signals from a frozen teacher, yet how a sampled loss allocates updates across tokens remains poorly understood. We analyze the gradient of the per-token K2 estimator of reverse KL with respect to the student logits. The $\ell_1$...
Bing Shao, Jia-Zheng Zhang, Long Ma et al.· 0 citations
Compared with typical vision-language tasks, document parsing places stronger demands on fine-grained visual perception. Existing vision-language model (VLM)-based parsing approaches rely on globally compressed visual tokens, where fine-grained details are entangled within a single representation and repeatedly accesse...
Ming-Xu Chai, Chen-Yu Liu, Zi-Yu Shen et al.· 0 citations
Reasoning-enhanced large language models have achieved remarkable improvements in planning tasks, yet their deployment in embodied systems remains impractical due to prohibitive inference delays-often exceeding minutes per planning instance. The fundamental bottleneck stems from the serial nature of existing paradigms:...
Yuchen Huang, Xijiang Ying, Zhenhua Ma et al.· 0 citations
This work adapt structured expert judgment from decision theory, using context-aware calibration questions to estimate expert reliability based on the quality of its probabilistic predictions, and employs Cooke-style log weighting, which penalises overconfident incorrect predictions and favours well-calibrated experts.
The Prefix-Adaptive Block Diffusion Model (PA-BDM) is proposed, which replaces intra-block bidirectional denoising with causal denoising from prefix to suffix and treats the block size as a maximum candidate range rather than a fixed commitment unit.
Ming-Xu Chai, Zi-Yu Shen, Chen-Yu Liu et al.· arXiv.org· 0 citations
CAFE (Coupled Agent--Feedback Evolution), a framework in which a shared-parameter model alternates between search-agent and critic roles, is introduced, suggesting that a self-improving search agent needs feedback that co-evolves with the policy it guides.
Bo-Yang Liu, Senjie Jin, Pei-Xin Wang et al.· 1 citation
AgentGym2 is presented, a new evaluation framework with task instances grounded in real-world end-to-end working demands that measures agents'ability to execute end-to-end procedures, discover tools via exploration, compose tools for unseen tasks, and remain robust to noisy and underspecified information.
Zhiheng Xi, Dingwen Yang, Jiaqi Liu et al.· Annual Meeting of the Associ...· 1 citation
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.