AutoDavis is introduced, a first-of-its-kind automatic and dynamic evaluation protocol that enables on-demand benchmarking of LVLMs across specific capability dimensions and shows effectiveness and reliability, offering a new paradigm for dynamic benchmarking of multimodal intelligence.
Han Bao, Yue Huang, Yan-Bo Wang et al.· Proceedings of the 32nd ACM...· 0 citations
This work establishes a question-level audit under fixed budgets, temperatures, and answer formats, and asks why reachable answers sometimes fail to appear, and test whether inference-time layer routing can expand reachability.
Yan-Chao Li, Wan-Hao Liu, Jia-Qing Xie et al.· 0 citations
AutoDavis is introduced, a first-of-its-kind automatic and dynamic evaluation protocol that enables on-demand benchmarking of LVLMs across specific capability dimensions and shows effectiveness and reliability, offering a new paradigm for dynamic benchmarking of multimodal intelligence.
Han Bao, Yue Huang, Yanbo Wang et al.· Proceedings of the 32nd ACM...· 0 citations
MASS learns low-dimensional principal manifold coordinates with a dense autoencoder for coarse semantic grouping, and then performs quality-aware sparse feature coverage within each group using a TopK sparse autoencoder and proposes MASS.
Peng Sun, Yi Yang, Antong Zhang et al.· 0 citations
Data-DPO, a target model-oriented SFT data selection method that consistently outperforms existing data selection baselines under multiple data budgets and stably surpasses full data training performance is proposed.
Peng Sun, Yi Yang, Antong Zhang et al.· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.