Visual token compression lowers the inference cost of vision--language models by representing images with fewer tokens. However, most existing methods compress visual tokens to a reduced set, leaving the amount of visual evidence represented by each token and its original spatial context implicit. Therefore, the compre...
Qi Zhang, Xian-Dong Meng, Rong-Gang Wang et al.· 0 citations
Protein design is moving beyond structural correctness toward function-aware design, yet existing generative models typically treat dynamics as a downstream property estimated through simulation or prediction after structure generation. Using MD trajectories as a generative target is also undesirable because stochastic...
Yu-Tian Liu, Mu-Jie Lin, Lan-Qian Zhang et al.· 0 citations
This work proposes CodecArena, the first vision-language framework for video coding quality assessment, casting codec evaluation as source-conditioned comparative reasoning between a reference and its reconstructions, and achieves state-of-the-art agreement with human judgments on source-disjoint content across diverse...
Jiaye Fu, Wei-Qi Li, Qiankun Gao et al.· 0 citations