Skip to content

Author

Yuanxing Zhang

We have 11 of 96 papers

We haven’t gathered this author’s papers yet. Follow them and we’ll fetch their work.

Not the right person? Other researchers publish under this name.

Preprint Sep 2026

Think Before You Score: Thinking Reward Model for Visual Generation

Visual reward models are essential for evaluating and improving visual generation models, yet existing approaches typically map task conditions and candidate outputs directly to scalar rewards, leaving implicit what should be evaluated for each individual case. We introduce Think Before You Score, a paradigm that expli...

Xue-Yuan Bai, Zhen-Chen Tang, Yang Shi et al. · 0 citations
Preprint Sep 2026

Human-Centric Image Captioning with Subject-Centered Spatial Understanding

While multimodal large language models (MLLMs) achieve remarkable performance on generic image captioning, they frequently suffer from structural hallucinations in human-centric scenarios. Accurately modeling human subjects is foundational for critical downstream applications, such as accurate avatar/video/image genera...

Bozhou Li, Jia-Hang Zhang, Yue Ding et al. · 0 citations
Jul 2026

Flux-OPD: On-Policy Distillation with Evolving Contexts

Flux-OPD is proposed, an OPD paradigm that uses evolving contexts as in-training supervision to capture task preferences in open-ended domains and outperforms existing OPD paradigms, highlighting the potential to combine teacher supervision with evolving contexts.

Yu-Ran Wang, Zekun Wang, Bohan Zeng et al. · 1 citation
#artificial intelligence Preprint Sep 2026

From Evaluation to Enhancement: Benchmarking and Improving Think-with-Video Reasoning for Video Generative Models

Video generation has advanced to produce visually compelling and temporally coherent results. Yet, whether these models can genuinely think with video--executing symbolic rules, respecting physical laws, and pursuing intentional goals--remains an open question. Existing benchmarks only partially address this, often con...

Meng Luo, Yi-Chen Liu, Jia-Hao Wang et al. · 0 citations
#artificial intelligence Preprint Sep 2026

FlowCPO: A Unified Divergence View of Preference Alignment for Flow Models

Preference alignment for flow and diffusion models now spans online reinforcement learning and offline preference optimization, but the relation between these methods remains unclear. In particular, existing forward-process alignment methods require fresh samples from the current model, while offline methods based on f...

Yan-Sen Han, Shengyi Liao, Peng Sun et al. · 0 citations
Preprint Jul 2026

RefCaptioner: Multi-Reference Image-Grounded Video Captioning

Existing video captioning models generate natural descriptions of video content but cannot explicitly ground local visual elements to multiple reference images. We introduce multi-reference image-grounded video captioning, a new task requiring factual video descriptions with phrase-level reference grounding, and propos...

Tengfei Liu, Yang Shi, Yuran Wang et al. · 0 citations
Jul 2026

Beacon: Knowing When and How to Perform Agentic Visual Reasoning

The fundamental goal of agentic visual reasoning is to improve the success rate of multimodal large language models (MLLMs) on complex tasks, rather than merely equipping them with a sophisticated yet inefficient reasoning paradigm. In this work, we rethink agentic visual reasoning through two key dimensions of tool us...

Qixun Wang, Yang Shi, Le-Tian Cheng et al. · 0 citations
Jul 2026

MultiRef-Compass: Towards Comprehensive Evaluation of Multi-Reference-to-Audio-Video Generation

MultiRef-Compass is introduced, a unified benchmark for MR2AV generation that integrates automatic metrics with a rejudging-enhanced MLLM-as-a-Judge framework, enabling scalable and auditable evaluation of both perceptual fidelity and reference-conditioned composition.

Xiaohan Zhang, Yu-Qing Wen, Jun-Lin Chen et al. · 3 citations
Jul 2026

ReBind: Multi-Reference Video Editing via Structured Instructions with Explicit Reference Relationships

This work proposes ReBind, a systematic framework that introduces semantic instructions with embedded reference tokens as the intermediate representation for multi-reference image-conditioned video editing and develops ReBind-Instruct, a specialized MLLM that learns to establish explicit bindings between visual attribu...

Xin-Yu Liu, Shi-Hao Li, Weihong Lin et al. · 2 citations
Jul 2026

AVE-Compass: Towards Holistic Evaluation for Audio-Video Editing Abilities

AVE-Agent is proposed, a modular agent framework that decomposes complex instructions into dependent subtasks and iteratively improves editing results through self-reflection and evaluator feedback, and improves instruction execution, Fidelity Preserving, and audio-visual alignment in joint editing while maintaining co...

Yuqing Wen, Yu-Kai Huang, Qianqian Xie et al. · 0 citations

MM-BrowseComp: A Comprehensive Benchmark for Multimodal Browsing Agents

The introduction of MM-BrowseComp, a novel benchmark comprising 400 challenging, hand-crafted questions designed to evaluate multimodal retrieval and reasoning capabilities, is introduced, establishing MM-BrowseComp as a rigorous new standard for the field.

Shilong Li, Xingyuan Bu, Wenjie Wang et al. · 37 citations · ⚡7

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.