Large Language Model (LLM) pretraining performance is jointly shaped by three components of the training triplet: the optimizer, model architecture, and training data stream. However, how these components influence performance in distinct ways remains unclear. We take a first step toward isolating their effects by stud...
Feng-Zhuo Zhang, Shu-Che Wang, Sheng-Gui Li et al.· 0 citations
This work presents WorldReward, a VLM-based pairwise preference reward model that unifies action-consistency and visual-quality evaluation for camera-conditioned world models, and introduces WorldReward-Bench, a human-annotated benchmark measuring reward-model agreement with human preferences across action consistency,...
Yi-Bin Wang, Ze-Han Wang, Junshu Tang et al.· 0 citations
The predictive divergence mask is proposed, which asks whether the next policy-gradient step will increase or decrease the same divergence used by the trust region, and the resulting masks improve RL training across model scales and precision settings.
JarvisHub is introduced, a canvas-native creative agent harness for long-horizon multimodal creation, where agents can progressively plan, generate, revise, and organize multimodal projects while users remain able to inspect, guide, and intervene throughout the process.