Educational foundation models must solve problems, understand curriculum structure, diagnose learner difficulties, and provide appropriate instructional support. Existing educational language models often focus on either problem solving or tutoring, with training mixtures organized by source or task rather than capabil...
Hao Liang, Qi-Han Lin, Mei-Yi Qiang et al.· 0 citations
Data policies for reinforcement learning with verifiable rewards (RLVR) determine which rollouts are used, how strongly they are weighted, and which domains contribute to subsequent training batches. We introduce DataFlex-RL, an evaluation platform for comparing these choices under a common GRPO recipe. Our primary exp...
Hao Liang, Mingrui Chen, Hengyi Feng et al.· 0 citations
WorkSurface-Bench is introduced, a benchmark for evaluating the capability of heterogeneous knowledge sources as surface routing in enterprise agents, and shows that correct surface selection is necessary but insufficient for task completion.
Hao Liang, Meiyi Qiang, Sizhe Qiu et al.· 0 citations
OmniaBench provides a broad and diagnostic benchmark for characterizing the capability boundaries of general agents across diverse scenarios with explicit state spaces, and introduces a ten-dimensional capability taxonomy and eight compositional atomic difficulty factors to support fine-grained evaluation and analysis.