This work proposes LLaVA-Assessor, a unified data construction and model training system for LMM-based machine vision, and introduces a simple yet effective prompt disentanglement strategy to alleviate training-objective confusion in multi-task learning, thereby enabling stable and coherent joint training.
Zi-Heng Jia, Zi-Cheng Zhang, Jia-Ying Qian et al.· 0 citations
TRWORLDBENCH is introduced, a benchmark for evaluating embodied world models through synchronized head, left-wrist, and right-wrist videos and uses 19 metrics to assess tri-view consistency, task alignment, physical and 3D coherence, motion quality, temporal consistency, and visual quality.
Xuan-Yi Liu, Hao-Feng Wang, Rui-Qi Li et al.· 0 citations
This paper introduces SafeAtlas-VL, a dataset of 1.5M training instances that places image-, request-, and response-level judgments on a five-level ordered scale, and trains the SafeAtlas Guard series of models via target-conditioned tuning for multimodal safety detection.
Zong-Rui Wang, Xiang-Yang Zhu, Sixiang Wang et al.· 0 citations
ELBench is introduced, the first benchmark to evaluate all four requirements (General Capability, Safety and Trustworthiness, Basic Education, and High-Level Cultivation) on the same models under a common protocol, combining curated public sources with newly synthesized safety and cultivation data.
Yi-Lin Jiang, Xiao-Rong Zhu, Fei Tan et al.· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.