SA-OPD is proposed, a Spurious-Signal-Aware On-Policy Distillation framework that identifies and filters misleading token-level supervision based on input-groundedness and optimization impact, and consistently outperforms Vanilla OPD and competitive selective methods.
Yinuo Jiang, Yongjie Ye, Zhou Tao et al.· 4 citations
This work identifies and formalizes the Distribution-Value Coevolution principle: the training value of data is not intrinsic, but emerges dynamically from the interaction between data characteristics and the model's evolving capability boundary, and operationalizes this principle through a unified framework.
Zairun Yang, Yanbo Yang, Chenyi Zhou et al.· Proceedings of the 32nd ACM...· 0 citations
SciToolAgent-Evo, an ontology-aware self-evolving agent for open-world scientific tool acquisition, driven by an evolving memory of skills, experiences, and an ontologized tool graph achieves state-of-the-art performance, validating its robustness and generalization.
Yuqi Tang, Chenyi Zhou, Libin Wang et al.· arXiv.org· 0 citations
EM^2Mem is proposed, an event-centric multimodal memory framework that binds heterogeneous evidence to event anchors during memory construction, enabling compact evidence readout over grounded multimodal events rather than modality-specific fragments.
Yijun Chen, Yangfan Zheng, Yanyang Li et al.· 0 citations
It is argued that effective knowledge editing must account for the intricate nature of knowledge representation, and three promising research directions are proposed that respect the complexity of knowledge representation in a real-world setting.
The same harness runs across five backend LLMs from three model families, indicating the harness generalizes across backends without tuning, even as different models induce distinct execution styles under the same workflow.
Jingsheng Zheng, Xinyuan Fang, Jintian Zhang et al.· 0 citations
Reinforcement learning from human feedback (RLHF) has become the cornerstone of aligning large language models (LLMs) with human intent. Yet a fundamental question remains unaddressed: how should training data be scheduled when both the model's capabilities and the utility of data are constantly evolving? Current pipel...
Zairun Yang, Yanbo Yang, Chenyi Zhou et al.· Proceedings of the 32nd ACM...· 0 citations
SciAtlas is presented, a shared, machine-actionable cross-disciplinary scholarly knowledge infrastructure that integrates evidential, conceptual, disciplinary, expertise, and normative layers under a shared schema and achieves a unified neuro-symbolic retrieval mechanism that grounds heterogeneous research objects, pro...
Shuofei Qiao, Yun-Xiang Wei, Bu-Sheng Zhang et al.· 1 citation
WorldMind is introduced, a framework that autonomously constructs a symbolic World Knowledge Repository by synthesizing environmental feedback that unifies Process Experience to enforce physical feasibility via prediction errors and Goal Experience to guide task optimality through successful trajectories.
Baochang Ren, Yunzhi Yao, Rui Sun et al.· arXiv.org· 3 citations· ⚡1
OceanGym is introduced, the first comprehensive benchmark for ocean underwater embodied agents, designed to advance AI in one of the most demanding real-world environments, and reveals substantial gaps between state-of-the-art MLLM-driven agents and human experts.
Yida Xue, Mingjun Mao, Xiangyuan Ru et al.· 0 citations
Mechanist is an agentic system that uses AI as a scientific instrument for the autonomous discovery of mechanisms underlying AI intelligence, and develops a mechanism theory of belief, revealing how models represent world knowledge, form beliefs, infer the beliefs of others, and how these mechanisms emerge during pretr...
Mengru Wang, Jun-Feng Fang, Shuo-Fei Qiao et al.· 1 citation
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.