Recent medical multimodal models have benefited from larger corpora, broader modality coverage, and stronger reasoning-oriented training, yet effective data design across continued pretraining (CPT) and post-training remains challenging. Medical sources vary substantially in structure, granularity, and information dens...
Guang-Hao Zhu, Ze-Yu Liu, Zhitian Hou et al.· 0 citations
SMAT (Simple MAT) is introduced, which jointly optimizes expert loss and expected loss at simulated merged parameters generated by sampling scaling coefficients, masks, and additive noise, and improves the mean score across five merging methods.
Yang-Gan Gu, Yuan-Yi Wang, Zhen Li et al.· 0 citations
Model merging efficiently combines specialized large language models (LLMs) without joint retraining, but can substantially alter expert routing in Mixture-of-Experts (MoE) models. Such \emph{routing drift} is often interpreted as routing failure, raising a fundamental question that remains unclear: \emph{does routing...
Yuan-Yi Wang, Yang-Gan Gu, Su Lu et al.· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.