Ensuring the safe and reliable deployment of large language models (LLMs) remains a fundamental challenge. Existing safety alignment approaches either incur high computational cost or unintentionally disrupt the model's core knowledge, leading to degraded fluency and factual accuracy on benign tasks. This reveals a per...
Ji-Sheng Dang, Yu-Shu Zhao, De-Wei Liu et al.· 0 citations
This work proposes Aesthetic Alignment: aligning generated images to explicit, user-specified compositional constraints using the Principles of Art (PoA), and introduces CompArt, a dataset of 80,032 WikiArt images augmented with captions and PoA analyses produced by a multimodal LLM under structured prompting.
Zheng-Hao Jin, Tat-Seng Chua· International Conference on...· 1 citation
The multi-view text-guided multimodal fusion adapter (MVFA) is proposed, a parameter-efficient framework that augments frozen LLMs with strong multimodal reasoning capability and achieves state-of-the-art performance on key metrics while updating only a small fraction of parameters.
Peng-Fei Shao, Ji-Sheng Dang, Jia-Wen Fang et al.· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.