UltraG-Bench: A Multi-task Benchmark for assessing Large Vision-Language Models on Pixel-level Evidence Grounding in Ultrasound
This work proposes UltraG-Agent, which combines the semantic reasoning capabilities of a VLM with the ultrasound-specific segmentation capability of UltraSAM3, and shows that UltraG-Agent substantially improves both semantic prediction and pixel-level visual grounding.
Quan-Hao Zhu, Bo Xu, Rui Lin et al.
· 0 citations