How often models reward-hack without instructions to do so, how effective and detectable their methods are when hacking is allowed, and how they adapt when an LLM review panel returns its decision and reasons are studied.
Yue Huang, Zhangchen Xu, Yu-Chen Ma et al.· 1 citation
Large language models (LLMs) are increasingly deployed as tool-augmented agents, where responses can depend on tool calls and external observations rather than model parameters alone. This creates an evaluation mismatch for LLM unlearning: previous unlearning methods may suppress direct parametric recall, but an agent...
Baicheng Chen, Zhe-Yuan Liu, Jingyu Zhang et al.· 0 citations
This work introduces a flow-matching framework with a linear interpolation path between paired view representations, that replaces diffusion with probability flows between observed and missing views, and provides a formal analysis showing that deterministic ODE flows are inherently better aligned with clustering object...
Yiteng Yuan, Junyan Wang, Zheyuan Liu et al.· arXiv.org· 0 citations
ManGo (Manga Active Narrative Grounding Optimization), an unsupervised framework for active manga visual question answering, introduces Active Narrative Sketching (ANS), which iteratively selects panels, extracts concise grounded clues, and decides when to stop, forming a compact question-directed evidence sketch befor...
Hao Qiu, Jun-Yan Wang, Zheyuan Liu et al.· 0 citations
A knowledge-verified benchmark that first confirms through a neutral probe that an agent knows a user's entitlement, and then evaluates whether it makes false claims once an incentive to deny that entitlement is introduced, which reduces the confound between lying and not knowing and enables more rigorous auditing and...
Zhe-Yuan Liu, Wei-Liang Zhao, Xiangchi Yuan et al.· 0 citations
We introduce UniWorld-Design, a framework that redefines image generation from flat pixel synthesis to structured visual composition, with semantic RGBA layers as the atomic units of generation, understanding, and editing. Our key insight is that pixels define how an image is rendered, whereas layers define how an imag...
Zongjian Li, Zhi-Yuan Yan, Chenxu Bai et al.· 0 citations
Safety-aware Contrastive Decoding (SafeCoDe) is introduced, a lightweight and model-agnostic decoding framework that dynamically adjusts token generation based on multimodal context that consistently improves context-sensitive refusal behaviors while preserving model helpfulness.