HarMoE: Multi-Source Chest Radiograph Pretraining with Dataset-Disentangled Experts
Recent vision-language models for chest X-ray understanding are largely built on image-report alignment and therefore rely heavily on MIMIC-CXR as the dominant pretraining source. While effective at scale, this paradigm underexplores an important alternative source of supervision: a range of existing multi-label classi...