Visual reward models are essential for evaluating and improving visual generation models, yet existing approaches typically map task conditions and candidate outputs directly to scalar rewards, leaving implicit what should be evaluated for each individual case. We introduce Think Before You Score, a paradigm that expli...
Xue-Yuan Bai, Zhen-Chen Tang, Yang Shi et al.· 0 citations
While multimodal large language models (MLLMs) achieve remarkable performance on generic image captioning, they frequently suffer from structural hallucinations in human-centric scenarios. Accurately modeling human subjects is foundational for critical downstream applications, such as accurate avatar/video/image genera...
Bozhou Li, Jia-Hang Zhang, Yue Ding et al.· 0 citations
Flux-OPD is proposed, an OPD paradigm that uses evolving contexts as in-training supervision to capture task preferences in open-ended domains and outperforms existing OPD paradigms, highlighting the potential to combine teacher supervision with evolving contexts.
Video generation has advanced to produce visually compelling and temporally coherent results. Yet, whether these models can genuinely think with video--executing symbolic rules, respecting physical laws, and pursuing intentional goals--remains an open question. Existing benchmarks only partially address this, often con...
Meng Luo, Yi-Chen Liu, Jia-Hao Wang et al.· 0 citations
Preference alignment for flow and diffusion models now spans online reinforcement learning and offline preference optimization, but the relation between these methods remains unclear. In particular, existing forward-process alignment methods require fresh samples from the current model, while offline methods based on f...
Yan-Sen Han, Shengyi Liao, Peng Sun et al.· 0 citations
Existing video captioning models generate natural descriptions of video content but cannot explicitly ground local visual elements to multiple reference images. We introduce multi-reference image-grounded video captioning, a new task requiring factual video descriptions with phrase-level reference grounding, and propos...
Tengfei Liu, Yang Shi, Yuran Wang et al.· 0 citations
The fundamental goal of agentic visual reasoning is to improve the success rate of multimodal large language models (MLLMs) on complex tasks, rather than merely equipping them with a sophisticated yet inefficient reasoning paradigm. In this work, we rethink agentic visual reasoning through two key dimensions of tool us...
Qixun Wang, Yang Shi, Le-Tian Cheng et al.· arXiv.org· 0 citations
MultiRef-Compass is introduced, a unified benchmark for MR2AV generation that integrates automatic metrics with a rejudging-enhanced MLLM-as-a-Judge framework, enabling scalable and auditable evaluation of both perceptual fidelity and reference-conditioned composition.
This work proposes ReBind, a systematic framework that introduces semantic instructions with embedded reference tokens as the intermediate representation for multi-reference image-conditioned video editing and develops ReBind-Instruct, a specialized MLLM that learns to establish explicit bindings between visual attribu...
Xin-Yu Liu, Shi-Hao Li, Weihong Lin et al.· arXiv.org· 2 citations
AVE-Agent is proposed, a modular agent framework that decomposes complex instructions into dependent subtasks and iteratively improves editing results through self-reflection and evaluator feedback, and improves instruction execution, Fidelity Preserving, and audio-visual alignment in joint editing while maintaining co...
The introduction of MM-BrowseComp, a novel benchmark comprising 400 challenging, hand-crafted questions designed to evaluate multimodal retrieval and reasoning capabilities, is introduced, establishing MM-BrowseComp as a rigorous new standard for the field.
Shilong Li, Xingyuan Bu, Wenjie Wang et al.· arXiv.org· 37 citations· ⚡7
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.