ESCD aggregates prefix-related teacher events and supervises the total probability of byte-compatible one-step student completions, avoiding tokenizer-dependent probability splits among individual tokens, and support event entry and event completion as complementary supervision targets for cross-tokenizer knowledge tra...
Jia-Cheng Liu, Jing-Wei Song, Qi-Tuan Zhang et al.· 0 citations
Recursive self-improvement (RSI) seeks to enable AI systems to participate in improving their own capabilities. A concrete pathway is autonomous model development, where agents iteratively explore post-training strategies to improve a base model. This setting faces two challenges: agents may exploit open-ended experime...
Ya-Xin Du, Xi-Yuan Yang, Zhi-Fan Zhou et al.· 0 citations
Answer-Backtracked Credit Assignment (ABC) is proposed, a fine-grained credit assignment framework for training long-horizon search agents by converting sparse trajectory-level outcomes into dense step-level supervision that rewards useful actions (even in failed trajectories) while suppressing erroneous or redundant a...
Yi-Jun Lu, Rui Ye, Jia-Jun Wang et al.· 1 citation
Learning graph structures from data is a fundamental problem that spans a wide range of signal processing and machine learning tasks. While significant effort has been made to tackle the problem, existing research has largely evolved along two parallel directions. The first seeks to infer the topology of an individual...
Xiaowen Dong, Hoi-To Wai, Si-Heng Chen et al.· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.