The Evaluation Agent framework is proposed, which employs human-like strategies for efficient, dynamic, multi-round evaluations, offering detailed, user-tailored analyses and is efficient, promptable, explainable, and scalable across models and tools.
Shu-Lin Tian, Zi-Qi Huang, Fan Zhang et al.· 2 citations
This work formalizes probabilistic alignment as a distributional criterion for world models and introduces PAWBench, a benchmark for evaluating video generators as stochastic samplers of world dynamics, and introduces PAWEval, an outcome-level protocol that converts repeated video rollouts into empirical distributions...
Yuandong Pu, Le Zhuo, Sayak Paul et al.· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.