Ensuring the safe and reliable deployment of large language models (LLMs) remains a fundamental challenge. Existing safety alignment approaches either incur high computational cost or unintentionally disrupt the model's core knowledge, leading to degraded fluency and factual accuracy on benign tasks. This reveals a per...
Ji-Sheng Dang, Yu-Shu Zhao, De-Wei Liu et al.· 0 citations
The multi-view text-guided multimodal fusion adapter (MVFA) is proposed, a parameter-efficient framework that augments frozen LLMs with strong multimodal reasoning capability and achieves state-of-the-art performance on key metrics while updating only a small fraction of parameters.
Peng-Fei Shao, Ji-Sheng Dang, Jia-Wen Fang et al.· 0 citations
The results support frozen verification as a training signal for evidence selection, while showing that strict boundary precision remains comparatively weaker.
Ming-Wen Zhang, Ji-Sheng Dang, Min-Qiang Yang et al.· 0 citations
PhysMLLMs is a training-stage prior injection architecture that injects physics-inspired spatial continuity priors into Video MLLMs, demonstrating that the injected spatial prior improves video consistency without compromising image-level grounding or general multimodal capability.
Siyao Yan, Bo Han, Ji-Sheng Dang et al.· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.