A Conditional Denoising Diffusion Probabilistic Framework for Synthetic Data Augmentation for Thoracic Disease Classification
Abstract
Thoracic disease localization and classification using deep learning models on chest X-rays are limited due to a.) imbalanced medical datasets, b.) severe overfitting, and c.) poor generalization on unseen data. This paper proposes a two-stage training pipeline based on a Conditional Denoising Diffusion Probabilistic Model (CDDPM) to generate high-fidelity synthetic images of chest X-rays in four diagnostic categories: Normal, Pneumonia, COVID-19, and Tuberculosis. The resulting synthetic corpus of 6,000 images is first $p$ re-trained on downstream classifiers and then fine-tuned on actual clinical data. We compare our proposed methodology with three state-of-the-art architectures VGG-16, EfficientNet-B0, and S win Transformer and observe that the classification accuracy, recall, and F1-score are substantially enhanced. It is important to note that VGG-16 has a 6.61% higher accuracy and Swin Transformer has the highest, with overall 92.61% accuracy. In order to apply this technique, we use Denoising Diffusion Implicit Model (DDIM) sampling which is 20 times faster than the standard DDPM inference. Our findings support the idea that diffusion-based synthetic augmentation is a promising approach to address data scarcity in medical imaging.