Preprint
Aug 2026
What Proves You Wrong: Benchmarking Language Models on Falsifiable Research Ideation
The Lit2Test benchmark centers on a six-field contract organized around a falsifying outcome, so that every proposal precommits the observation that would prove it wrong, making its quality decidable in the first place rather than merely arguable.
Ziyue Wang, Aomufei Yuan, Yiran Yao et al.
· 0 citations