The Evaluation Agent framework is proposed, which employs human-like strategies for efficient, dynamic, multi-round evaluations, offering detailed, user-tailored analyses and is efficient, promptable, explainable, and scalable across models and tools.
Shu-Lin Tian, Zi-Qi Huang, Fan Zhang et al.· 2 citations
Apple-PI is introduced, the first benchmark that anchors video-model evaluation explicitly in physical laws, and is positioned as a diagnostic foundation for guiding future video models toward world models with law-grounded physical intelligence.
Runmao Yao, Kairui Hu, Yukang Cao et al.· 1 citation
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.