In this research, we aim to evaluate the robustness of convolutional neural networks (CNNs) and foundation models, such as ConCH and UNI, in classifying different types of organs from whole-slide images (WSIs) collected from various countries, different scanners, and clinical environments. Despite the dataset's inherent diversity, our results reveal that even a small perturbation (a white-box attack) with an intensity of 0.01 significantly impacts performance. ResNet-50 experienced a 67% reduction, ConCH 19%, and UNI 9.5% in accuracy. However, the foundation models performed much better, showing greater resilience and maintaining comparatively strong predictive performance even under the same intensity. Still, randomizing the dataset does not necessarily make your model resilient or robust, emphasizing the importance of thorough generalization and stability testing before clinical deployment. These models should be rigorously tested to evaluate their generalization and reliability prior to application in real patient settings. Foundation models look promising for building reliable and general AI-based cancer diagnostic systems, but they still need to prove their reliability in real-world clinical environments before large-scale adoption.
Khan Ziaullah, Md Ariful Islam Mozumder, Hee-Cheol Kim· 2026 6th International Confe...· 0 citations
International Classification of Diseases (ICD) codes enable correct billing, insurance reimbursement, and healthcare analytics. However, manual coding is time-consuming, expensive, and error-prone, creating bottlenecks in clinical workflow and limiting scalability. Artificial intelligence (AI) has emerged as a promising solution for automated ICD code assignment from unstructured clinical text. This systematic review explores the current state of automated ICD coding research, examining models applied to diverse clinical documents including discharge summaries, electronic health records, nursing notes, and pathology reports. Following PRISMA guidelines, we searched six databases for studies published between 2019 and 2024, selecting 54 relevant studies from 4,280 initial citations. Our analysis reveals the use of diverse datasets, preprocessing techniques, and feature extraction methods, alongside a clear evolution from traditional machine learning to deep learning approaches, with substantial architectural diversity across convolutional, recurrent, transformer, and hybrid models. Performance varies considerably across dataset configurations, with models achieving higher accuracy on frequent code subsets compared to full label spaces. However, critical gaps persist: overreliance on single-language, single-institution datasets limits generalizability; difficulties in predicting rare codes remain unresolved; lack of model interpretability undermines clinical trust; and inconsistent evaluation protocols hinder meaningful comparison. To address these challenges, we propose a 5P evidence-grounded research agenda: Population Diversity, Performance Robustness, Prediction of Rare Codes, Provenance Transparency, and Practical Integration. These findings underscore AI’s potential to transform ICD coding while highlighting the need for standardized benchmarks, rigorous external validation, multilingual datasets, and explainable architectures to enable equitable and effective deployment in real-world healthcare systems.
Abdul Rehman Khalid, Haider Ali, Kounen Fathima et al.· Journal of medical systems· 0 citations
The imperative necessity for rapid discovery of antiviral agents against emerging viral diseases, such as COVID-19 caused by SARS-CoV-2, has emphasised the limitations of conventional drug discovery, with its deliberate pace, high costs, and high failure rate. In this study, we present a high-throughput “AI-driven virtual screening pipeline” that combines cutting-edge molecular embeddings using transformers for compounds and language model-based protein sequences for targets, and combines these using a gradient-boosted decision tree-based model (XGBoost) trained on ChEMBL bioactivity data, with excellent performance for binary prediction (ROC-AUC 0.8408, Accuracy 0.76) comparable to top-performing methods. When applied to the large library of natural products from COCONUT database, our model correctly predicted top-ranked compounds, some of which were shortlisted for molecular docking against the SARS-CoV-2 spike receptor-binding domain (RBD), the key interface for ACE2 binding, in order to test the validity of proposed pipeline and the results revealed several promising compounds with excellent binding energies and multiple interactions with hotspot residues such as K417, Y453, Q493, G496, Q498, N501, Y505 and F486. The potential of cost-effective, synergistic integration of scalable deep learning of representations, interpretable machine learning, and physics-based refinement of structure is significant for accelerating natural product-based therapeutics for coronaviruses and possible other viral threats, with the potential to expand to ensemble approaches, active learning, and variant-based targets for enhanced efficacy.
Hayat Ullah, Khan Ziaullah, Md Ariful Islam Mozumder et al.· 2026 6th International Confe...· 0 citations