This work proposes UltraG-Agent, which combines the semantic reasoning capabilities of a VLM with the ultrasound-specific segmentation capability of UltraSAM3, and shows that UltraG-Agent substantially improves both semantic prediction and pixel-level visual grounding.
Hateful memes spread abusive content through implicit interactions between images and text, posing serious threats to the safety of online communities. In recent years, multimodal large language models have been widely used for hateful meme detection and are increasingly adopted to generate explainable detection result...
Bo Xu, Chen-Yuan Wang, Xin-Yu Chen et al.· 0 citations
This work proposes MedVCoT, which incorporates latent visual reasoning into the medical visual question answering (VQA) domain, and utilizes the specialized expertise of MedSAM to train a large vision-language model so that it can autonomously generate consistent and continuous latent visual tokens within Visual Chain-...
Bo Xu, Quan-Hao Zhu, Bo-Lin Zhu et al.· Proceedings of the Thirty-Fi...· 1 citation
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.