The Graph Theory Agent (GTA), which pairs a preference-trained representation selector with plan-and-decompose scaffolding around a frozen executor LLM, is proposed, which lifts Phi-4 from 53.5% to 69.1% on the benchmark's easy split and from 33.0% to 41.5% on its hard split.
Zi-Xiang Xu, Yan-Bo Wang, Chenxi Wang et al.· 2 citations· ⚡1
SocialMaze is introduced, a benchmark that organizes six tasks across social deduction games, daily-life interactions, and digital community platforms along three descriptive design axes: deep reasoning, dynamic interaction, and information uncertainty.
Zi-Xiang Xu, Yan-Bo Wang, Yue Huang et al.· 1 citation
Existing studies of LLM-as-judge scoring bias work predominantly at the input-output level: they perturb inputs, measure score deltas, and propose prompt-level mitigations. We argue that the same biases admit a representation-level account in the judge's hidden state, complementary to the input-output view and operatio...
Zixiang Xu, Sixian Li, Hua-Xing Liu et al.· arXiv.org· 2 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.