Skip to content

Author

Peng-Fei Zhang

We have 4 of 13 papers

We haven’t gathered this author’s papers yet. Follow them and we’ll fetch their work.

Not the right person? Other researchers publish under this name.

Jul 2026

Listen, Do Not Copy: Internalizing Audio-Grounded Scaffold Context for Robust Omni-Model Speech Understanding

Omni models transcribe clean, single-speaker speech well, but their accuracy drops sharply when speakers overlap and the scene is noisy, exactly where knowing who said what matters most. A natural fix is a short scene description. We show why this is risky: answer-bearing text lets the model copy instead of listen, so...

Peng-Fei Zhang, Biao Tian, Tianxin Xie et al. · 0 citations
#machine learning Preprint Sep 2026

The Platonic brain bridge hypothesis: human brain networks as an architectural prior for multimodal large language models

Multimodal large language models predict brain activity, but brain alignment has been a measurement, not a design tool. We propose the Platonic brain bridge hypothesis: omni models, multimodal large language models that process video, audio and text jointly, converge on brain-like representations usable in both directi...

Peng-Fei Zhang, Biao Tian, Xian-Gang Li et al. · 0 citations
Preprint Aug 2026

DreamX-Phi 1.0: Action-Conditioned Video World Model for Robotic Manipulation

We present \textbf{DreamX-Phi 1.0}, an action-conditioned video world model for robotic manipulation that, given an observed frame, a language instruction, and a prescribed action sequence comprising end-effector poses and gripper states, predicts the resulting future observations. Yet realism alone does not guarantee...

Dream Team, Rui Chen, Xiangxiang Chu et al. · 3 citations
Jul 2026

Evaluation-Verification Reward for Consistent Multi-Reference Image Editing

A Multi-dimensional Evaluation-Verification Reward (EVR) decomposes evaluation into distinct visual criteria; for each criterion, an MLLM Evaluator generates multiple candidate hypotheses, and a Verifier grounds each claim in concrete visual evidence to accept or reject it, producing reliable and fine-grained reward si...

Ying-Mao Miao, Peng-Fei Zhang, Xiaochen Lv et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.