As LLM agents undertake increasingly complex, long-horizon tasks, verifying their outputs becomes increasingly challenging. We study how verification capability can be strengthened with a fixed base model, without access to reference answers or grading rubrics at test time. Repeated sampling yields multiple rollouts th...
Caiqi Zhang, Ru-Jun Han, Zifeng Wang et al.· 0 citations
This work proposes XConf (eXperiential Confidence): estimating confidence together with the model's accumulated experience, and sees experiential confidence estimation as a new paradigm for future general-purpose confidence estimation.
Caiqi Zhang, Xiao-Chen Zhu, Chengzu Li et al.· 0 citations
It is shown that raw global calibration metrics are not robust for cross-model comparison, and that fair calibration comparison requires accuracy-aware evaluation, and proposed ACE, an accuracy-controlled evaluation framework, is proposed.
Zhichao Yang, Caiqi Zhang, Ruihan Yang et al.· arXiv.org· 1 citation
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.