Frame selection is a fundamental component of multimodal large language models, enabling long videos to be processed under limited visual-token and computational budgets. Uniform sampling preserves temporal coverage but may miss informative content that appears only briefly. To alleviate this limitation, query-dependen...
Xinlei Pu, Weijie Shi, Wen Yang et al.· 0 citations
A new harness self-evolution method, named DREvo, is proposed, which integrates function-level evidence anchoring, state-dependent evidence recalibration, and role-conditioned search intent distillation to determine which historical evidence remains valid and where the harness should evolve next.
SkillReranker is proposed, an inference-time reranking framework for adaptive skill selection that effectively improves task performance, reduces environment interaction steps, and lowers token consumption compared with existing skill selection baselines.
Yanping Chen, Weijie Shi, Wen Yang et al.· arXiv.org· 3 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.