Skip to content

Author

T. Tegegne

2 papers indexed here

We haven’t gathered this author’s papers yet. Follow them and we’ll fetch their work.

Not the right person? Other researchers publish under this name.

Review Jul 2026

Artificial intelligence for predicting antiretroviral therapy outcomes in people living with HIV: A systematic review of predictive models, predictors and clinical readiness.

BACKGROUND Machine learning (ML), deep learning (DL) and other predictive modelling approaches are increasingly applied to predict antiretroviral therapy (ART) outcomes among people living with HIV, yet their methodological robustness and suitability for clinical and digital health integration remain uncertain. OBJECTIVE To systematically review artificial intelligence (AI)-based models predicting ART outcomes among people living with HIV, focusing on predictors, methodological rigour, validation practices and clinical integration readiness. METHODS For this review, AI was considered an umbrella concept encompassing ML and deep learning methods, while conventional predictive models such as logistic regression were synthesized separately were included in comparative studies. This review followed PRISMA 2020 guidelines and was registered in PROSPERO (CRD420251166548). Seven databases were searched from January 2000 to 25 September 2025. Data were extracted on model types, predictors, outcomes, validation strategies and performance metrics. Methodological quality and reporting completeness were assessed using PROBAST and the TRIPOD Statement. We proposed an exploratory clinical informatics framework for evaluating the translational readiness across data, workflow, interoperability and operational domains. This framework was applied to assess characteristics relevant to clinical readiness. RESULTS Twenty-six studies were included. Virologic outcomes (34.6%), retention in care (30.8%) and adherence (19.2%) were most common. Ensemble methods, particularly Random Forest (65.4%) and XGBoost (23.1%), predominated, while deep learning approaches were less frequent (15.4%). Key predictors included duration on ART (61.5%), CD4 count (69.2%), adherence history (57.7%), age (73.1%) and sex (65.4%). Reported discrimination varied widely (area under the curve [AUC] 0.56-0.99), but interpretation is limited by poor calibration, limited external validation and potential overfitting. Most studies showed a high or unclear risk of bias. Validation practices were limited: 80.8% relied solely on internal validation, 19.2% performed external validation and 30.8% reported calibration assessment. Reporting completeness was also suboptimal, with frequent omissions of confidence intervals and model specifications, indicating incomplete adherence to TRIPOD. No study fulfilled all clinical informatics domains, highlighting a persistent gap between model development and real-world implementation. Longitudinal modelling was limited (15.4%), and random data splitting in such datasets introduced a risk of data leakage. CONCLUSIONS Although AI models demonstrate promising predictive performance, their clinical applicability is limited by substantial methodological limitations, including high risk of bias, inadequate validation, poor calibration and limited transparency. Strengthening external and temporal validation, routine calibration, TRIPOD-compliant reporting and integration into clinical workflows will be essential to support reliable deployment, particularly in resource-limited settings.

B. Dejene, Yaregal Assabie, Mulugeta Tadele et al. · 0 citations
Open access Jul 2026

An explainable AfroXLMR approach for multi-label emotion classification of Amharic social media text with dataset release.

Emotion detection from social media is crucial for understanding human emotions across languages. However, for low-resourced languages such as Amharic, the lack of annotated data makes this task challenging. Additionally, most current models use black-box methods that obscure whether predictions rely on linguistically meaningful cues. To address these gaps, this study proposes a multi-label emotion classification model for Amharic by fine-tuning AfroXLMR. To enhance transparency, we integrate explainable artificial intelligence (XAI) into the framework. We compiled and annotated a new dataset of 22,000 unique social media comments across eight emotion categories for training, validation, and testing. The data was split into 80% for training, 10% for validation, and 10% for testing. The proposed model achieved a recall of 87% and a Hamming loss of 0.08. To interpret its predictions, we applied Local Interpretable Model-agnostic Explanations (LIME). We also evaluated the model against several state-of-the-art baselines, including XLM-R base, mBART, BiLSTM, LSTM, CNN, and AfriBERTa. The results show that our approach outperformed each baseline, achieving F1-score improvements of 5% over XLM-R base, 3% over mBART, 5% over BiLSTM, 7% over LSTM, 9% over CNN, and 2% over AfriBERTa. Bootstrapped statistical significance testing confirms that these improvements are robust and not attributable to random variation. In conclusion, the fine-tuned AfroXLMR model demonstrates promising performance in Amharic multi-label emotion classification. Building on this success, next steps could involve exploring more advanced fine-tuning strategies and expanding our datasets to strengthen both performance and the model's ability to generalize across diverse Amharic contexts.

Yeshimebet Bayu, Demeke Endalie, T. Tegegne · 0 citations