While Reinforcement Learning (RL) effectively incentivizes reasoning in Large Language Models, current pipelines are hindered by training instability and rapid entropy collapse. These limitations often stem from"Rollout Silencing"and low-quality gradient signals in standard sampling procedures. In this work, we propose...
Yi-Meng Ye, Shuang Chen, Wen-Xuan Huang et al.· 0 citations
JarvisHub is introduced, a canvas-native creative agent harness for long-horizon multimodal creation, where agents can progressively plan, generate, revise, and organize multimodal projects while users remain able to inspect, guide, and intervene throughout the process.
Video-DR is introduced, featuring a decoupled perception-exploration pipeline with stage-wise tool unlocking that compels exhaustive cross-frame visual grounding prior to web retrieval, enabling autonomous exploration that breaks the imitation-learning ceiling.