LigBench is proposed, an automated evaluation benchmark that enables fine-grained and reliable evaluation of AI research ideas, consistently applicable across different generation distributions, and PAIR-IQ, a dataset tailored for training pairwise idea judgment models and serving as an auxiliary reference to support more objective comparative evaluation.
Chenrun Wang, Mingxuan Zhu, Tiancheng Huang et al.· 0 citations
MASS learns low-dimensional principal manifold coordinates with a dense autoencoder for coarse semantic grouping, and then performs quality-aware sparse feature coverage within each group using a TopK sparse autoencoder and proposes MASS.
Peng Sun, Yi Yang, Antong Zhang et al.· 0 citations
Data-DPO, a target model-oriented SFT data selection method that consistently outperforms existing data selection baselines under multiple data budgets and stably surpasses full data training performance is proposed.
Peng Sun, Yi Yang, Antong Zhang et al.· 0 citations