A benchmark should deliver more than a scalar score: what makes an evaluation trustworthy is the reasoning that justifies the score. This is especially critical for world models, where judging a rollout requires understanding whether physics, causality, and world state evolve correctly. Humans spot such violations natu...
Weiliang Chen, Haowen Sun, Jun Gao et al.· 2 citations
Post-training has been shown to significantly improve language models'performance on tasks with verifiable outcomes, including mathematical reasoning, software engineering, and computer use. However, whether the same approach can improve forecasting in financial markets is much less clear. Compared with tasks with veri...
Jia-Cheng Guo, Suo-Zhi Huang, Shu-Zhen Li et al.· 0 citations
Knowledge Graph Construction (KGC) is essential for transforming unstructured text into structured knowledge representations. Despite advances in Large Language Models, existing methods treat KGC as a single-pass generation task, conflating extraction, normalization, and validation within a single forward pass. This le...
Shi-Long Liu, Yongqiang Liu, Jiye Liu et al.· Proceedings of the 32nd ACM...· 0 citations
This work designs a contrastive study that combines controlled quantitative experiments with paired trajectory analysis and consolidates observations into a taxonomy of three high-level categories and twelve skill-use modes, showing that skills work when noisy trajectories become procedural anchors that stabilize execu...
Zhi-Yuan Jiang, Fan Huang, Hanwen Xing et al.· 2 citations
Extensive experiments demonstrate that AgentsKG outperforms state-of-the-art training-free baselines in both extraction accuracy and structural quality, offering a robust approach to open-domain knowledge graph construction without additional training.
Shilong Liu, Yongqiang Liu, Jiye Liu et al.· Proceedings of the 32nd ACM...· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.