Predicting item difficulty from content can provide an initial estimate for newly developed questions before sufficient student responses are available. Existing approaches typically represent the question stem and answer choices as text. When mathematics items contain visual components, a common pipeline first textual...
Generative and agentic AI are reshaping both the production and evaluation of scientific research. These developments are often studied separately, as questions of how AI can produce research and how AI can review it. We argue that this separation misses an increasingly important feature of scholarly publishing: change...
Chen-Guang Wang, Ming Li, A. Braimah et al.· 1 citation
This work introduces A 2 -Judger, a novel MLLM-based A gentic instantiation of A uto Judger equipped with semantic-aware retrieval and dynamic memory that significantly improves sample efficiency while maintaining reliable evaluation results.
Xuanwen Ding, Chengjun Pan, Zejun Li et al.· 0 citations
An exam-style evaluation framework is introduced for studying the global budget allocation of reasoning language models when multiple problems share an end-to-end cost or latency constraint, in which a model must distribute one shared token budget across questions with different difficulty and point values.
Chenrui Fan, Yize Cheng, Ming Li et al.· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.