LastOPD is proposed, which applies the latent signal only at the last-layer state, the common interface both LM heads read, and only during a 10-step crossfade into token-level OPD, which keeps the useful part of the latent signal and hands the student to token-level supervision before the collapse sets in.
Jie Yang, Zheng-Yu Fang, Ze-Lin Xu et al.· 2 citations· ⚡1
TimeEvo is proposed, which clusters an agent's diagnosed failures into capability gaps, plans a measurement for each, synthesizes evidence-only tools that fill them, and admits the candidate library only through a paired admission gate.
Jie Yang, Yan Zheng, Jia-Rui Sun et al.· 0 citations
The evaluation of large language models (LLMs) on coding tasks has primarily focused on performance metrics such as pass@k. As LLMs continue to advance, many models now meet baseline performance requirements, reducing the discriminative power of performance-based evaluation alone. Yet a key question remains largely une...
Jun-Peng Wang, Yuzhong Chen, Menghai Pan et al.· 0 citations
Bazaar is introduced, a dynamic sealed-bid benchmark for multi-attribute auction under multi-attribute auction under these conditions, grounded in closed-form customer utilities, enabling exact evaluation.
Shimaa Ahmed, Yiwei Cai, Mohsen Minaei et al.· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.