Educational foundation models must solve problems, understand curriculum structure, diagnose learner difficulties, and provide appropriate instructional support. Existing educational language models often focus on either problem solving or tutoring, with training mixtures organized by source or task rather than capabil...
Hao Liang, Qi-Han Lin, Mei-Yi Qiang et al.· 0 citations
CapGeo-Bench is proposed, a benchmark of 4,641 high-quality figure-caption pairs equipped with a fine-grained keypoint-based evaluation metric that provides high-quality captions consistently and substantially boosts performance of MLLMs, empirically validating the visual perception bottleneck in geometric reasoning.
Yu-Ying Li, Si-Yi Qian, Hao Liang et al.· 5 citations
Repo2Skill-Evo casts each release transition as a skill-maintenance task: given a V1 skill set and the official V1-to-V2 patch, an agent must update obsolete skill content while preserving guidance that remains valid.
Chenyuan Duan, Ge Shi, Zineng Mao et al.· 0 citations
WorkSurface-Bench is introduced, a benchmark for evaluating the capability of heterogeneous knowledge sources as surface routing in enterprise agents, and shows that correct surface selection is necessary but insufficient for task completion.
Hao Liang, Meiyi Qiang, Sizhe Qiu et al.· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.