Leveraging the paradigm of autoregressive diffusion, text-conditioned interactive video world models aim to simulate temporally coherent environments guided by textual instructions. While enabling low-latency, long-duration generation is pivotal for embodied AI and simulation-based planning, current frameworks primaril...
Ze-Yu Zhang, Jin-Yuan Mao, Da-Kai An et al.· 0 citations
Multimodal large language models (MLLMs) process hundreds or thousands of visual tokens per image, incurring prohibitive inference costs. While existing vision token pruning methods mitigate this overhead, they implicitly assume that a single fixed pruning strategy can be applied uniformly across all inputs. Our analys...
Hai-Jin Liang, P. Zhou, Zheng-Lin Wan et al.· 0 citations
This work proposes Reinforcement Learning with Human-Engine Verification (RLHEV), a post-training paradigm that combines dense engine signals with implicit human acceptance feedback from the development process to support RL post-training.
P. Zhou, Hesong Wang, Zhengfeiyang Zhang et al.· 0 citations
This work proposes SmartRes, a framework that performs efficiency optimization in the pixel space via dynamic resolution routing and introduces a margin-regularized routing objective that increases foreground-background logit separation and improves foreground recall.
H. Sun, Wang-Bo Zhao, Fanyue Wei et al.· 0 citations
Reinforcement Learning with Verifiable Rewards (RLVR) makes Multimodal Large Language Models more accurate, but the gains are brittle: simply paraphrasing a question or changing the prompt template can degrade them, which challenges reliable deployment in high-stakes scenarios like medical VQA. We trace this to two iss...
P. Zhou, Zhiwei Tang, Xiaopeng Peng et al.· 0 citations
Spiking neural networks (SNNs) offer a path to energy-efficient language modeling through sparse encoding and event-driven computation, but training capable spiking language models from scratch remains difficult. A practical alternative is ANN-to-SNN migration through knowledge distillation (KD), where a pretrained art...
En-Qiao Lu, Xingrui Yu, Yi-Wei Fu et al.· 0 citations
CRM+RCCR, an architecture-agnostic cost-aware objective that encodes cost preference into continuous relevance targets through per-pair independent scoring, eliminating multi-positive dilution while regularizing queries with similar routing preferences to be closer in the routing space.
Tao Yu, Yi-Fei Qu, Zhi-Qing Cui et al.· 1 citation
By overcoming the longstanding memory and scalability barriers, RPG serves as a critical advance in ‘ AI generating AI ’, potentially enabling efficient weight generation at scales previously deemed infeasible.
Kai Wang, Dongwen Tang, Wangbo Zhao et al.· Neural Information Processin...· 7 citations· ⚡1
This work introduces UniHall, a fine-grained dataset grounded in a unified taxonomy spanning Object, Instruction, and Knowledge dimensions, and proposes Self-Adaptive Multimodal Fuzzing (SAMF), a self-adaptive framework that employs evolutionary mutation strategies to explore the boundaries of model hallucinations.
P. Zhou, Jiajun Song, Zhiwei Tang et al.· 1 citation
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.