Lingshu: Generalist Foundation Model for Unified Multimodal Medical Understanding and Reasoning.
Multimodal Large Language Models (MLLMs) excel at understanding generic visual content, such as landscapes, objects, and events, thanks to extensive datasets and advanced training regimes. However, their effectiveness in medical applications remains limited due to the inherent discrepancies between data and tasks in me...