Sparse autoencoders (SAEs) are proposed to extract numerous features from large language model (LLM) representations, yet explaining these features still relies primarily on external observation. This reliance leads to superficial explanations inferred from observed model behavior and computational inefficiency from co...
Weihang Meng, Hongzhu Guo, Yi Jing et al.· 0 citations
Supervised Fine-Tuning (SFT) and Reinforcement Learning (RL) exhibit fundamentally different behaviors in enhancing multi-task reasoning for large language models (LLMs). Our preliminary experiments revealed a phenomenon: SFT suffers from severe task conflicts under multi-stage training, whereas RL enables stable coexi...
Kejian Zhu, Zhuo-Ran Jin, Shangqing Tu et al.· 0 citations
This work proposes SurveyReview, a reviewer-aligned, multi-dimensional benchmark and dataset for survey evaluation, and develops a strong baseline evaluator that substantially improves alignment with human reviewers, providing a competitive reference for future research.
Yuheng Zhang, Yuanchun Wang, Fanjin Zhang et al.· Proceedings of the 32nd ACM...· 0 citations
RuVerBench is introduced, the first benchmark for assessing LaaJ reliability in rubric verification for agentic scenarios, and the impact of key LaaJ strategies, including prompt design, batching, and majority voting, on rubric verification is analyzed.
The rapid advancement of large language models has transformed survey writing from a months-long manual effort into an automated process. As generation scales, reliable evaluation becomes the bottleneck, and LLMs are increasingly used as survey evaluators. However, existing approaches largely rely on off-the-shelf LLM-...
Yuheng Zhang, Yuanchun Wang, Fanjin Zhang et al.· Proceedings of the 32nd ACM...· 0 citations
Tail subtraction is introduced, which removes shared prompt and continuation semantics from boundary states and yields cleaner, more stable steering signals, and suggests that steering depends on representations of what the model is about to do, not merely on what has already appeared.
Jiaran Ye, Lingxu Ran, Zijun Yao et al.· arXiv.org· 2 citations
TrajDebug is proposed, an error-lifecycle tracing framework that addresses long-trajectory error discovery with multi-granularity history compression and evidence-based error identification, and supports critical attribution by tracing each error's resolution status and terminal impact.
Yunjia Qi, Zehua Yin, Xin Shi et al.· 4 citations· ⚡1
DeepWeaver is a novel framework that weaves noisy retrieved evidence into comprehensive answers by maintaining Thought Block Chains (TBCs), a structured representation that groups claims, salient information, keywords, and supporting evidence.
Xujia Wang, Yizheng Zhang, Bin Xu et al.· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.