Research Attention Prediction (RAP), a rolling benchmark covering 278 AI/ML fields and 1,390 episodes, is introduced, a rolling benchmark covering 278 AI/ML fields and 1,390 episodes for large language models to track shifts in research attention.
Ying-Qian Wu, Jingcong Liang, Si-Yuan Wang et al.· 0 citations
This survey provides a first systematic overview of the emerging area of vision meets graphs, which treats visual depictions of graphs as first-class inputs for reasoning and learning, and organize existing work into three threads.
Xinjian Zhao, Wei Pang, Zhixuan Yu et al.· Proceedings of the Thirty-Fi...· 2 citations
The CVPR 2026@AdvML Workshop Challenge on adversarial multimodal attacks against autonomous-driving VLAs is presented, providing a practical reference for future robustness evaluation and defense design in multimodal autonomous-driving systems.
Tian-Yuan Zhang, Zonglei Jing, Jiangfan Liu et al.· arXiv.org· 0 citations
MatrAIx is introduced, a population-scale simulated-user evaluation infrastructure for testing AI systems and digital products with heterogeneous users and provides an end-to-end infrastructure for evaluating AI systems and digital products with diverse simulated human users.
Xiaomin Li, Yuexing Hao, Jian Hou et al.· 1 citation
Self-Supervised Visual On-Policy Distillation (S$^2$VOPD), a simple yet effective method that constructs on-policy learning signals from asymmetric augmented views, systematically explores a broad design space of visual augmentations and uncover that asymmetry matters.
Yijiang Li, Yijun Liang, Yunjie Tian et al.· 0 citations
SciOrch is presented, a framework that trains a lightweight 8B model to orchestrate frontier LLMs for scientific reasoning, and attains the best accuracy on both SGI and SFE with less than half the API cost of typical multi-agent methods.
Jingru Guo, Xiangyuan Xue, Lian Zhang et al.· arXiv.org· 0 citations
It is found that configuration shift consistently erodes CP validity, often driving empirical coverage below the target, and coverage lower bounds are derived that attribute this loss to a discrepancy between calibration and test score distributions.
Yuqicheng Zhu, Jia-Lin Yu, Lin Li et al.· 0 citations
This work introduces VBVR-Pro, a closed-loop testbed that makes native visual reasoning through generation trainable, verifiable, optimizable, and experimentally controllable, and identifies recurring failure modes of the prevalent VLM-as-a-judge paradigm.
Junhua Xu, Rui-Si Wang, Fanyi Pu et al.· 5 citations
This work develops a multiagent simulation of a popular social network, Reddit, and uses millions of posts from users on the platform to model content-sharing on the platform.
Swapneel Mehta, Bogdan State, Richard Bonneau et al.· 1 citation
Progress in large language models is often summarized using a single scalar measure, such as a time horizon, a latent ability estimate, or an aggregate benchmark score. These summaries capture the overall performance, but they do not test whether progress is distributed differently across task difficulty. We find that...
Hanwen Xing, Pengyu Wang, Bingxu Meng et al.· 1 citation
This work constructs and releases SETA-Env, the largest open-source verifiable terminal RL dataset to date, containing over 4,500 environments, and demonstrates that SETA- Env provides high-quality training environments for terminal agents and serves as a valuable resource for advancing research on terminal-based agent...
Q. Shen, Zhiqi Huang, V. Kamanuru et al.· 8 citations· ⚡2
This study provides empirical clarity through a systematic exploration of multimodal pretraining and derives efficient pretraining recipes that achieve strong generative performance using only 5% of the compute budget.
Junlin Han, Shengbang Tong, David Fan et al.· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.