Financial search is a highly demanding task for LLM agents, requiring not only a correct final answer but also temporally valid information retrieval, authoritative source selection, entity and period alignment, unit and definition consistency, and verifiable evidence for all conclusions. Existing benchmarks predominan...
Wen-Qing Wang, Hai-Tao Xiang, Xin-Yi Zhao et al.· 0 citations
Experiments across diverse real-world web-generation scenarios show that LiveEvalBench aligns closely with human expert judgment and provides fine-grained insights into frontier models' web generation capabilities.
Yiyao Wang, Zhen Wen, Ying Tang et al.· 0 citations
It is observed that OPD-trained models maintain superior avg@K performance across sampling budgets, while the advantage in pass@K gradually shifts to the pre-OPD base models as K increases, suggesting that OPD primarily improves sampling efficiency rather than consistently expanding the student's reasoning capability b...
Xinmu Ge, Zizhuo Zhang, Yu Huang et al.· 1 citation
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.