Existing auto-research benchmarks often entangle multiple sources of improvement, including training frameworks, hyperparameters, compute budgets, and data, making it difficult to attribute why one frontier agent outperforms another to specific research capabilities. In this work, we isolate and systematically evaluate...
Rui-Feng Yuan, Yi-Zhi Li, Ya-Xin Du et al.· 0 citations
Multimodal Large Language Models (MLLMs) excel at understanding generic visual content, such as landscapes, objects, and events, thanks to extensive datasets and advanced training regimes. However, their effectiveness in medical applications remains limited due to the inherent discrepancies between data and tasks in me...
Wei-Wen Xu, Hou-Pong Chan, Long Li et al.· IEEE Transactions on Pattern...· 0 citations
This work investigates the fundamental design principles required for robust visual text representation learning and trains Pixel Linguist II, a native-resolution vision encoder trained with on-the-fly rendering, unified contrastive grounding, and a multilingual curriculum over 280M training examples.
Chaohao Yuan, Rui-Feng Yuan, Zhuoxu Huang et al.· 2 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.