Cloud-deployable deep learning for smartphone-based body fluid cytology: a multi-architecture evaluation of YOLO and RT-DETR models
Abstract
Accurate cytological evaluation of body fluid specimens is essential for diagnosing inflammatory and malignant conditions, yet manual interpretation remains labor-intensive and subject to inter-observer variability. This study developed and deployed a smartphone-based deep learning framework for multi-class body fluid cell classification integrating object detection, ROI-based classification, and explainable artificial intelligence (XAI). Three hundred Wright–Giemsa–stained cytospin slides, one per patient, were used to generate 7149 smartphone-acquired images, which were annotated and curated into seven predominant cell classes. To reduce data leakage, images were partitioned at the source-photograph (slide-field) level. Five detection architectures (YOLOv8n, YOLOv10n, YOLOv11n, YOLOv11n-seg, and RT-DETR) were trained and evaluated on a held-out test set. YOLOv11n-seg achieved the highest mAP@50 (0.840), mAP@50–95 (0.757), and precision (0.772), while YOLOv11n achieved the highest recall (0.823). For YOLOv11n-seg, confusion-matrix macro-F1 and weighted-F1 were 0.720 and 0.730, respectively. Eosinophils, lymphocytes, and neutrophils were among the better-recognized classes, whereas errors were more frequent among monocytes, macrophages, and mesothelial cells. Grad-CAM++ analysis showed that model activation was predominantly centered within cell regions but was interpreted only as supportive model-inspection evidence. A Streamlit web application demonstrated browser-based inference with detection overlays, cell counts, differential composition, and data export. The system should currently be regarded as research, educational, and decision-support prototype rather than a standalone diagnostic device. External validation with patient-level separation, multiple imaging configurations, and broader cell populations is required before clinical application can be considered.