Skip to content
Conference

PharyTriFuse: Knowledge- and LLM-Augmented Deep Learning for Bacterial Pharyngitis Detection from Smartphone Throat Images

Jul 2026 · Annual International Computer Software and Applications Conference · pp. 2290-2295 · 0 citations · 20 references

Abstract

Bacterial pharyngitis requires timely antibiotic treatment, whereas most non-bacterial cases are self-limited; diagnostic errors may therefore lead to missed infections or unnecessary antibiotic use. This study proposes PharyTriFuse, a multimodal framework that integrates throat-image analysis with large language model (LLM) reasoning and a medical knowledge graph (KG) to classify bacterial versus non-bacterial pharyngitis from smartphone-acquired oropharyngeal images. Experiments were conducted on the public PGUPharyngitis dataset over 742 images using a stratified 72%/8%/20% train/validation/test split. Images were standardized using CLAHE and redness enhancement to reduce acquisition variability. Two visual backbones (EfficientNet-B4 and ConvNeXt-Base) were evaluated under four configurations: AI-only, AI+LLM, AI+KG, and AI+LLM+KG. Performance was assessed using standard classification metrics and inference efficiency. Results show that incorporating LLM reasoning and structured medical knowledge improves classification performance over vision-only baselines while maintaining real-time inference capability under certain configurations. These findings suggest that multimodal AI systems can enhance smartphone-based decision support for pharyngitis assessment.

View source

Similar papers

#machine learning Preprint Aug 2026

AppendiGrade: An XAI-Enhanced Deep Learning Framework for Grading Appendicitis in Ultrasound with Gaussian Blur and Grad-CAM

An advanced system capable of automatically detecting complicated appendicitis from ultrasound images was developed and was explained with gradient-weighted class activation mapping (Grad-CAM), which creates a heatmap of the regions responsible for the model's prediction of the infected areas.

Fahad Ahammed, Omar Faruq Shikdar, Navid Zaman et al. · 0 citations
Open access Jul 2026

DEEP TRANSFER LEARNING FOR DERMOSCOPIC SKIN LESION CLASSIFICATION: BENCHMARKING XCEPTION, INCEPTIONRESNETV2, MOBILENETV3LARGE, DENSENET121, AND NASNETMOBILE

This research proves that transfer learning, systematic class balancing, and focal loss functions provide a computationally viable approach and highly effective method for automatic skin cancer classification, while also highlighting that the backbone architecture is the key factor that determines the effectiveness of classification under similar training conditions.

Yusra Shafiq, H. M. Shahzad · 0 citations
Review Open access Jul 2026

Uncertainty-Aware Prediction Across Endoscopic Domains: Laryngeal Narrow-Band and Gastrointestinal Imaging

Background: Deep-learning systems for endoscopic image classification are commonly evaluated with random data splits, which may overestimate performance under acquisition shift; uncertainty-aware selective prediction may improve reliability by allowing a model to abstain on uncertain cases. Methods: We evaluated binary abnormality detection in two endoscopic imaging domains: laryngeal contact-endoscopy narrow-band imaging (CE-NBI; 210 patients, patient-level) and gastrointestinal endoscopy (HyperKvasir; 6746 images, image-level, predominantly white-light). ImageNet-pretrained ResNet-50 deep ensembles were assessed under a resolution-defined acquisition-shift stress test; for the gastrointestinal data, a random stratified split was additionally used as an in-distribution reference. We evaluated discrimination, calibration, decision-curve analysis, and entropy-based selective prediction. Results: In the gastrointestinal dataset, random-split evaluation produced high performance (AUROC 0.994, 95% CI 0.990–0.996; AUPRC 0.991). Under resolution shift on the same data, performance fell to AUROC 0.723 (95% CI 0.710–0.736; AUPRC 0.690; sensitivity 0.417; specificity 0.874); the two intervals do not overlap. Selective prediction improved reliability among retained cases: under resolution shift, accuracy rose from 0.683 at full coverage to 0.852 (95% CI 0.835–0.871) at 25% coverage (balanced accuracy 0.646 → 0.798). Predictive entropy was significantly higher for incorrect than for correct predictions in both regimes (Mann–Whitney p = 7.1 × 10−77 with rank-biserial |r| = 0.30 under shift). In the laryngeal cohort, no statistically significant differences were detected among four architectures (ROC-AUC 0.844–0.901; all pairwise DeLong p > 0.05). Conclusions: Random-split evaluation substantially overestimated performance relative to a resolution-defined acquisition-shift stress test, and entropy-based selective prediction improved reliability by identifying a high-confidence subset for automated prediction while deferring the remainder to human review. Target-domain recalibration substantially restores calibration under shift (ECE 0.172 → 0.036 with temperature scaling; → 0.017 with isotonic regression) but does not recover discrimination; selective prediction is complementary, mitigating residual confident-wrong predictions. An encoder-transfer experiment showed asymmetric cross-domain utility; features learned on the larger gastrointestinal cohort transferred to the laryngeal cohort (AUROC 0.80 vs. in-domain 0.89), whereas the reverse direction did not transfer (0.53 vs. 0.72). Prospective multi-center validation remains required before clinical deployment. All code, fold definitions, random seeds, and a reproducible protocol are publicly released.

Behnam Kiani Kalejahi, S. Khan, Murodbek Akhrorov et al. · 0 citations
Open access Jul 2026

Development of a deep learning model for intussusception using point-of-care ultrasound

The feasibility of developing a deep learning model for the detection of intussusception using a smaller dataset of POCUS images is demonstrated and fine-tuning models were best adapted to the screening nature of POCUS images.

A. Thyagachandran, Brian Lefchak, H. Murthy et al. · 0 citations
Open access Jul 2026

AI-driven diagnosis of mpox using deep learning models

These grouped original-only results are intentionally conservative relative to augmentation-heavy or single-split designs and should be interpreted as deflated but more trustworthy reference values, and should be interpreted as a reproducible reference benchmark rather than a clinically validated diagnostic tool.

Bassam W. Aboshosha, Shafiq Ul Rehman, L. N. Mahmoud et al. · 1 citation