Deep Learning–Based Instance Segmentation for Automated Dental Diagnosis: A Comparative Study of Mask R-CNN, YOLOv8, Faster R-CNN, and DETR Architectures
Abstract
This study presents a comparative evaluation of deep learning architectures for automated instance segmentation of teeth and dental pathologies, including cavities, caries, and cracks, in dental images. Four state-of-the-art models—Mask R-CNN with ResNet50-FPN-V2 backbone, YOLOv8m, Faster R-CNN, and DETR—were implemented and trained on a standardized COCO-format dental dataset. The Mask R-CNN model was developed using PyTorch and Torchvision with optimized hyperparameters, including an initial learning rate of 0.001, ReduceLROnPlateau scheduling, and early stopping to prevent overfitting. The results of our experiments show that Mask R-CNN has better accuracy (86.1%), precision (87.2%), recall (85.0%), and F1-score (85.9%) than both YOLOv8m (67.51%), Faster R-CNN (63.8%), and DETR (31.69%). The results suggest that two stage detectors utilizing pixel level mask supervision more adequately address the task of instance segmentation for dental applications than either YOLO, Faster R-CNN, or DETR. This framework shows promise for use in CAD systems as it will increase the reliability of diagnostics and improve the clinical efficiency of them.