Long-horizon search requires agents to gather evidence across multiple steps and synthesize it into well-supported answers. The recent agent harnesses provide a natural and promising framework to support such long-running search processes. As interaction histories grow, one single agent in harnesses might get stuck and...
Shan-Yong Wang, Zhen-Wen Ji, Lei Jin et al.· 0 citations
Detailed image captioning requires accurate and comprehensive descriptions of fine-grained visual content, yet caption quality spans factual accuracy, information coverage, and clarity. Compared with conventional methods that rely mainly on high-quality supervision or holistic rewards, rubric-based reinforcement learni...
Zhen-Wen Ji, Lei Jin, Shanyong Wang et al.· 0 citations
A training-free, Adaptive Visual Token Pruning (AVTP) framework, applicable to diverse LVLM architectures is proposed, and adaptive pruning ratios in multi-image contexts where images of higher importance retain proportionally more tokens are implemented.
Rongyang Zhang, Cheng-Qiang Lu, Cong Li et al.· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.