As LLM agents undertake increasingly complex, long-horizon tasks, verifying their outputs becomes increasingly challenging. We study how verification capability can be strengthened with a fixed base model, without access to reference answers or grading rubrics at test time. Repeated sampling yields multiple rollouts th...
Caiqi Zhang, Ru-Jun Han, Zifeng Wang et al.· 0 citations
Conditional Progressive Pruning (CPP) is proposed, a lightweight pruning framework that fully exploits multi-round MAD and is the first to fully outperform consistency methods.
Ruo-Song Ye, Caiqi Zhang, Jia-Hao Li et al.· 0 citations
This work proposes XConf (eXperiential Confidence): estimating confidence together with the model's accumulated experience, and sees experiential confidence estimation as a new paradigm for future general-purpose confidence estimation.
Caiqi Zhang, Xiao-Chen Zhu, Chengzu Li et al.· 0 citations
It is shown that raw global calibration metrics are not robust for cross-model comparison, and that fair calibration comparison requires accuracy-aware evaluation, and proposed ACE, an accuracy-controlled evaluation framework, is proposed.
Zhichao Yang, Caiqi Zhang, Ruihan Yang et al.· arXiv.org· 1 citation
This work conducts the first large-scale, systematic studies of multilingual calibration across six model families and over 100 languages, revealing that non-English languages suffer from systematically worse calibration.
Ej Zhou, Caiqi Zhang, Tiancheng Hu et al.· arXiv.org· 10 citations· ⚡1
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.