Visual tool use has emerged as a fundamental capability for multimodal agents to actively acquire evidence beyond a fixed image encoding. The prevailing recipe learns this capability from teacher-generated trajectories filtered for answer correctness, implicitly assuming that every successful demonstration provides eff...
Changhao Xiang, Shi-Lin Zhang, Zheng Ma et al.· 1 citation
Evaluation is shifting from static QA toward agentic settings where models act through external tools. We identify a critical yet underexplored capability within this space - dexterous visual tool use: fine-grained, closed-loop parameterized visual action in which models infer tool parameters from visual evidence, and...
Shudong Liu, Dong-Yang Chen, En-Ci Zhang et al.· 0 citations
VERA (Visual Evidence-Retaining strategy for long-horizon Agents), a training-free context manager built on deterministic rendering with no exposed memory operations, supporting a modality-preserving view of long-horizon context management.
Jiang-Feng Su, Cong Pang, Jiawei Hong et al.· 1 citation
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.