R-Select, a robust and scalable framework that optimizes data selection with 30 distinct quality metrics, introduces a novel hierarchical optimization strategy that consistently outperforms both heuristic baselines and model-based methods, offering a robust solution for high-quality data curation.
Xin Gao, Xiao-Yang Wang, Yun Zhu et al.· Proceedings of the 32nd ACM...· 0 citations
In recent years, online Direct Alignment from Preferences (DAP) has emerged as a popular alternative for Reinforcement Learning from Human Feedback (RLHF) due to its training stability and simplicity. In online DAP, training relies on preference data, each composed of a question and a pair of large language model (LLM)...
Chi Zhang, Jia-Chen T. Wang, Kun He et al.· Proceedings of the VLDB Endo...· 0 citations
The transition from architecture-centric scaling to data-centric refinement has established high-quality data as a critical determinant of Large Language Model performance, particularly for complex reasoning and instruction following. However, effective data selection remains a persistent bottleneck: simple heuristic f...
Xin Gao, Xiaoyang Wang, Yun Zhu et al.· Proceedings of the 32nd ACM...· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.