This work introduces PlanGuard, the first pre-execution detector that evaluates the physical safety of a complete multi-step plan in its current environment, and proposes Strong-Teacher Adaptive Compensation for On-Policy Distillation (STAC-OPD), which provides compact models with adaptive strong-teacher supervision al...
Jun-Chi Chen, Chang-Tao Miao, Yu Xiang et al.· 0 citations
This work proposes EraseSAE, a novel framework that leverages sparse autoencoders to achieve surgical concept erasure in DiT-based T2V diffusion models via a principled decompose-attribute-erase pipeline, and introduces the Partitioned Convolutional Sparse Autoencoder.
Xing-Hao Wang, Dong Li, Wei Yu et al.· 0 citations
VisCo is a training-efficient self-compression framework that reuses the pretrained VLM itself as an intrinsic compressor that compresses visual information using a small set of memory tokens and transfers hierarchical information from encoding to decoding.
Yupeng Zheng, Kai Zou, Bin Liu et al.· arXiv.org· 0 citations
A training-free framework that localizes each utterance's source span through Source Span Localization and transfers the associated audio-visual latent content from the source interval to the specified target interval through Region-Aware Latent Remapping is proposed.
Chao Zhou, Yilu Chen, Qi Chu et al.· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.