This work proposes Modality-Adaptive Decoding (MAD), a training-free method that adaptively weights modality-specific decoding branches based on task requirements based on task requirements, demonstrating that explicit modality awareness through self-assessment is crucial for robust multimodal reasoning.
Sangyun Chung, Se Yeon Kim, Youngchae Chee et al.· arXiv.org· 3 citations· ⚡2
Remote sensing multimodal large language models (RS-MLLMs) have advanced scene understanding and visual question answering over satellite imagery, yet localizing specific objects or changed regions remains challenging. Existing approaches rely on generating bounding box coordinates as token sequences, which is fragile...
J. Chung, Sungjune Park, Yeongyun Kim et al.· 0 citations
This work decomposes video into view-invariant and view-variant latents and recompose them under language supervision encouraging clean decomposition of the two streams, and proposes PRISM, that achieves state-of-the-art results on EgoExo4D, EgoExoLearn, AE2, even surpassing in-domain models under zero-shot setting.
Youngchae Chee, Ho-Su Lee, Sungjune Park et al.· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.