Deep Multiple Instance Learning for Pulmonary Embolism Diagnosis in SPECT Images
Abstract
Objective: In the computer-aided diagnosis of pulmonary embolism using three-dimensional ventilation/perfusion single-photon emission computed tomography (V/P SPECT) images, conventional globally supervised learning models exhibit limited sensitivity to local lesions due to the small volume of perfusion defects and the absence of detailed slice-level annotations. This study proposes a pseudo-label-enhanced Transformer multiple instance learning model to improve the accuracy of slice-level analysis and the overall diagnostic sensitivity. Methods: The proposed approach formulates each three-dimensional V/P SPECT image as a bag and the corresponding two-dimensional slices as instances. The model utilizes a two-dimensional Transformer as the foundational feature extractor and incorporates a feature pyramid structure to capture multi-scale dependencies. To mitigate optimization difficulties, the method generates instance-level pseudo-labels using ventilation-perfusion difference images. These pseudo-labels are combined with examination-level ground truth labels to construct a hybrid loss function for constrained model training. Results: On the test set, the model achieves an accuracy, sensitivity, specificity, F1 score, and AUC of 0.969, 0.964, 0.976, 0.972, and 0.994, respectively. These metrics outperform those of conventional convolutional neural network feature extractors and single pooling aggregation strategies (p < 0.05). The diagnostic performance of the model surpasses that of radiologists with three years of experience, achieving zero missed diagnoses and zero misdiagnoses in nine complex cases. With the assistance of the model confidence scores, the diagnostic accuracy of the radiologists improves from 0.906 to 0.979. Conclusion: The Transformer multiple instance learning model incorporating a feature pyramid and pseudo-label constraints effectively identifies subtle local lesions in three-dimensional medical images. The approach demonstrates high diagnostic performance and clinical utility in scenarios lacking dense annotations.