Language model agents are increasingly deployed in open-world tool environments, which require balancing exploring unknown capabilities and exploiting known ones. Existing methods face a performance-efficiency tradeoff: they either rigidly decouple exploration and execution or interleave them without coordination. We a...
Shengbin Yue, Hongru Wang, Siyuan Wang et al.· 0 citations
Practical video editing is not only pixel generation: an editor must turn a brief, a clip pool, music metadata, and hard constraints into an executable timeline. We study this decision layer as \emph{executable video-editing planning} and introduce RefineCut, which, unlike workflow systems that wrap a prompted frontier...
Hao-Yu Wang, Cheng Feng, Liuyang Bian et al.· 1 citation
VISA (Visual Instruction Synthesis Agent), an agentic framework that reformulates multimodal instruction synthesis as a self-evolving loop, consistently improves multimodal instruction following over strong baselines, while preserving general multimodal capability across seven public benchmarks.
Min Zeng, Guanxin Tan, L. Cen et al.· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.