Preprint
Aug 2026
LigBench: A Unified and Human-Aligned Benchmark for LLM-based Research Idea Generation
LigBench is proposed, an automated evaluation benchmark that enables fine-grained and reliable evaluation of AI research ideas, consistently applicable across different generation distributions, and PAIR-IQ, a dataset tailored for training pairwise idea judgment models and serving as an auxiliary reference to support more objective comparative evaluation.
Chenrun Wang, Mingxuan Zhu, Tiancheng Huang et al.
· 0 citations