Preprint
Aug 2026
From Simple QA to Deep Research: A Verifiable Benchmark Constructed through Iterative Task Evolution
This work introduces a verifiable benchmark of 500 deep research tasks spanning 31 topics and 10 major categories, with three query forms designed to probe complementary capabilities required for deep research.
Can Wang, Haoran Chen, Hao Gao et al.
· 0 citations