Aug 2026· Frontiers in Medicine· Vol 13· 0 citations· 34 references
TL;DR
Combining interpretable relapse prediction with reinforcement learning offers a personalized decision-support framework for optimizing herbal treatment intensity during UC remission maintenance.
Abstract
Maintaining remission in ulcerative colitis (UC) with traditional Chinese herbal medicine is challenging because fixed-intensity prescriptions may not accommodate changes in inflammatory burden. This study aimed to predict 52-week relapse and develop a data-driven policy for sequential adjustment of herbal treatment intensity.
We analysed a development cohort of 1,181 patients receiving herbal maintenance therapy and an independent external validation cohort of 96 patients. Twenty-one baseline clinical, biomarker, herbal-dose and behavioral variables, expanded to 31 features after one-hot encoding, were used to train and compare seven machine-learning models through stratified five-fold cross-validation and external validation. Discrimination, calibration and decision-curve net benefit were evaluated. Remission maintenance was subsequently modelled as a sequential decision problem using 4,141 scheduled-visit transitions. A gradient-boosted reward model was used for model-based off-policy evaluation of a net-clinical-benefit policy that penalized unnecessary treatment intensification in clinically quiescent patients.
Calibrated logistic regression provided the best overall performance, with an internal cross-validation area under the receiver operating characteristic curve (AUROC) of 0.783 (95% confidence interval [CI], 0.755–0.809) and an external AUROC of 0.710 (95% CI, 0.600–0.820). The corresponding Brier scores were 0.180 (95% CI, 0.169–0.191) and 0.215 (95% CI, 0.170–0.263), respectively. Medication adherence was the strongest predictor of relapse, followed by fecal calprotectin, C-reactive protein and Coptidis Rhizoma dose. The learned policy increased herbal intensity with increasing fecal calprotectin and achieved a higher estimated mean reward than observed clinician behavior (0.552 ± 0.018 versus 0.063 ± 0.021). Policy-concordant visits were associated with a higher next-visit remission rate than discordant visits (70.0% versus 59.6%).
Combining interpretable relapse prediction with reinforcement learning offers a personalized decision-support framework for optimizing herbal treatment intensity during UC remission maintenance. Because the policy was derived and evaluated using observational data and model-based off-policy methods, prospective clinical evaluation is required before routine implementation.
Objective Hypertension is a common yet frequently underdiagnosed comorbidity in psoriasis patients. Early identification and blood pressure control are critical to improving outcomes. Although machine learning (ML) is widely used in disease prediction, a model for hypertension risk within the psoriasis population remains unavailable. This study aims to develop and validate such a model in patients with psoriasis. Methods In this retrospective study, 2,957 psoriasis patients from a single tertiary center were used for model development and internal validation, and 567 psoriasis participants from the National Health and Nutrition Examination Survey (NHANES) served as the external validation cohort. After missForest imputation and consensus feature selection, the Synthetic Minority Oversampling Technique (SMOTE) was applied to the training set only. Nine machine learning algorithms were trained and evaluated for discrimination, calibration, and clinical utility. Shapley Additive Explanations (SHAP) were used for model interpretation. Results Nine nonredundant predictors were retained. The SMOTE-enhanced logistic regression model showed the most balanced and generalizable performance, with area under the receiver operating characteristic curve values of 0.850, 0.816, and 0.789 in the training, internal validation, and external validation cohorts, respectively, together with acceptable calibration and favorable net clinical benefit. SHAP identified age, dyslipidemia, and type 2 diabetes mellitus as the leading contributors. The final model was deployed as a publicly accessible web application. Conclusions This interpretable and externally validated machine learning model provides a practical tool for hypertension risk stratification in psoriasis patients and may support earlier identification and individualized preventive management in clinical practice.
Guo-Hua Xue, Xiaowu Guo, Jia-Qi Chen et al.· Digital Health· 0 citations
Aim To develop and validate a multimodal artificial intelligence (AI)-based prediction model for platinum-resistant recurrence in ovarian cancer by integrating clinical data, medical imaging, and medical knowledge resources, with the goal of improving early risk stratification and supporting individualized treatment decisions. This exploratory proof-of-concept study aims to assess the feasibility of multimodal fusion for this task; no external validation has been performed. Methods This study collected multimodal data from ovarian cancer patients, including clinical records from 214 patients treated at Renji Hospital affiliated to Shanghai Jiao Tong University School of Medicine between June 2020 and January 2025, imaging data from 218 patients comprising 5,053 CT and MRI images, and 1,000 high-quality medical literature sources published between 2023 and 2025. Patients were classified into a platinum-resistant recurrence group (PROC, n = 87, 40.7%) and a non-platinum-resistant recurrence group (NPROC, n = 127, 59.3%) according to whether recurrence occurred within 6 months after the last platinum-based chemotherapy. The platinum-free interval (PFI) was used only for outcome definition, not as a predictor. A multimodal prediction framework based on a Mixture of Experts (MoE) architecture was constructed, incorporating a clinical expert model, an imaging expert model, and a medical knowledge expert model. Model performance was evaluated using accuracy, recall, F1 score, and the area under the receiver operating characteristic curve (AUC) with 95% confidence intervals (bootstrap with 1,000 iterations, stratified by cross-validation folds), calibration (Brier score, Hosmer-Lemeshow test, calibration curve), and decision curve analysis (DCA). Results The proposed multimodal model demonstrated excellent predictive performance for platinum-resistant recurrence, achieving an accuracy of 0.98 (95% CI: 0.96–1.00), a recall of 0.95 (95% CI: 0.92–0.98), an F1 score of 0.98 (95% CI: 0.97–0.99), and an AUC of 0.96 (95% CI: 0.95–0.97). The model showed good calibration with a Brier score of 0.042 (95% CI: 0.031–0.058) and a Hosmer-Lemeshow test p-value of 0.31 (χ2 = 11.8, df = 10), indicating no statistically significant lack of fit. These results were superior to those of single-expert and conventional benchmark models. In the clinical expert evaluation, the model achieved an accuracy of 0.83, a recall of 0.81, an F1 score of 0.82, and an AUC of 0.83, showing competitive performance compared with random forest, support vector machine, gradient boosting machine, and Transformer-based models. Conclusion This exploratory study demonstrates that multimodal AI integrating clinical, imaging, and knowledge graph data can achieve strong internal predictive performance for platinum-resistant recurrence of ovarian cancer in a single-center retrospective cohort. However, the model is preliminary, has not been externally validated, and is not ready for clinical use. Independent multicenter validation is required before any clinical translation can be considered. This article should be viewed as a hypothesis-generating tool and a methodological proof-of-concept only.
Xuanxuan Zhao, Ling Ma, Peiquan Li et al.· Frontiers in Medicine· 0 citations
Background Acute postoperative protein depletion, including hypoalbuminaemia and hypoproteinaemia, frequently complicates colon cancer surgery and exacerbates adverse outcomes, yet early risk stratification remains challenging. Aims To develop a predictive model utilising a tabular foundation model for acute postoperative protein depletion in colon cancer patients, alongside an interpretable clinical web tool. Methods We retrospectively evaluated perioperative data from 812 colon cancer patients treated between 2020 and 2025. Following recursive feature elimination, eight traditional machine learning algorithms and the TabICLv2 tabular foundation model were trained. Discrimination, calibration, incremental risk stratification, and clinical utility were assessed using the area under the receiver operating characteristic curve (AUC), paired DeLong tests, calibration curves, Brier score, net reclassification improvement (NRI), integrated discrimination improvement (IDI), and decision curve analysis. SHapley Additive exPlanations (SHAP) were used to visualise feature contributions. Results TabICLv2 achieved the numerically highest validation AUC among the evaluated algorithms using nine selected predictors, with an AUC of 0.766 (95% CI: 0.699–0.832) and a Brier score of 0.158. DeLong tests showed no statistically significant AUC differences between TabICLv2 and XGBoost or Random Forest. NRI/IDI analyses indicated improved event reclassification across comparator models, with significant total NRI and IDI improvement over Random Forest, whereas incremental improvement over Logistic Regression and XGBoost was limited. Sensitivity analyses using albumin-only and total-protein-only endpoints showed broadly consistent discrimination. SHAP analysis revealed age, prealbumin (PA), and globulin (GLO) as the leading contributors to model predictions. A web-based calculator was subsequently deployed to facilitate clinical translation. Conclusion: TabICLv2 integrates demographic, nutritional, and immunological profiles to predict acute postoperative protein depletion with moderate discrimination. The accompanying application may provide adjunctive individualised risk assessment, but prospective multicentre validation is required before routine clinical implementation.
Xinke Cao, Linrui Han, Xinquan Zan et al.· Frontiers in Oncology· 0 citations
Objectives To evaluate inflammatory-nutritional indices in relation to breast cancer (BC) risk and mortality and develop a cross-ethnically validated prediction model. Methods From National Health and Nutrition Examination Survey (NHANES) 2005–2018, 485 BC patients and 16,838 female controls were included, with mortality follow-up through 2019. Weighted multivariate logistic and Cox regression assessed associations between seven inflammatory indices, two composite indicators, and BC risk/mortality. Multiple machine learning (ML) algorithms, including XGBoost, were used to construct risk models. The model was externally validated (NHANES other periods:1999-2004) and cross-ethnic validated. We prospectively enrolled Chinese treatment-naïve breast cancer patients and matched healthy controls for external validation. Results In fully adjusted models, the Advanced Lung Cancer Inflammation Index (ALI) was inversely associated with BC risk and all-cause mortality (highest vs. lowest tertile: odds ratio [OR] 0.64, 95% CI 0.45–0.91; hazard ratio [HR] 0.41, 95% CI 0.18–0.90). Conversely, neutrophil percentage-to-albumin ratio (NPAR), systemic inflammation response index (SIRI), and neutrophil-to-lymphocyte ratio (NLR) showed positive associations. ALI outperformed other indices in predicting mortality. XGBoost identified NPAR as the top predictive feature; the model incorporating inflammatory indices and age achieved an AUC of 0.832 on the test set, and a web-based dynamic nomogram incorporating these factors was developed. External validation yielded AUCs of 0.781 (NHANES) and 0.730 (Chinese cohort). Conclusions ALI (protective) and NPAR/SIRI/NLR (detrimental) are robust predictors of BC risk and mortality. The ML model demonstrates good predictive performance, but cross-ethnic validation highlights the need for population-specific calibration, which indicated the potential of ML approaches leveraging inflammatory-nutritional indices to enhance BC risk stratification and inform clinical decision-making.
Yue Li, Ting Ding, Xiaoyan Zhou et al.· Frontiers in Immunology· 0 citations
The use of glucocorticoids in sepsis remains controversial due to heterogeneous treatment responses across patient populations. More individualized treatment strategies are needed to guide glucocorticoid administration.
An offline reinforcement learning (RL) model was developed using the MIMIC-IV database and externally validated in the eICU-CRD. Adult patients with sepsis were included, and clinical trajectories were constructed using 24-hour time steps. Least absolute shrinkage and selection operator (LASSO) regression was applied for feature selection, and missing data were handled using multiple imputation. The problem was formulated as a Markov decision process, with glucocorticoid dosing defined as discrete actions. A conservative Q-learning (CQL) algorithm was used to learn treatment policies. Policy performance was evaluated using fitted Q-evaluation (FQE), and comparisons were made with clinician strategies. SHAP analysis was performed to interpret model decisions.
A total of 3,070 patients from the MIMIC-IV cohort and 372 from the eICUCRD cohort were included. The CQL policy achieved higher expected returns than clinician policy in both datasets. Higher expected returns were associated with higher survival rates. The model recommended a more selective glucocorticoid use pattern, with reduced overall usage and delayed initiation in some patients.
The RL model identified individualized glucocorticoid treatment strategies and demonstrated improved policy performance compared with clinician practice, suggesting its potential as a clinical decision support tool.
Yu-Jing Zhang, Lei Liang, Xin-xin Zhang et al.· Frontiers in Cellular and In...· 0 citations
This study aimed to develop and interpret a machine learning model for predicting postoperative recurrence of anal fistula using routine laboratory indicators and inflammation-related indices.
A total of 2,214 patients who underwent fistulectomy were included. Patients from wards 5, 11, 12, 13 and 14 (
n
= 1,772) were divided by stratified random sampling according to recurrence status into training (
n
= 1,242) and testing (
n
= 530) cohorts. Patients from wards 15 and 16 (
n
= 442), which were managed by separate clinical teams, were reserved as a ward-based internal validation cohort. Univariate and multiple fistula tracts. analyses were performed to identify recurrence-associated factors, and LASSO regression was used for feature selection. Multiple machine learning models were developed and compared, including logistic regression, support vector machine, GBM, neural network, XGBoost, AdaBoost, LightGBM, and CatBoost. Model performance was assessed using ROC curves, calibration curves, decision curve analysis, and classification metrics. SHAP analysis was applied for model interpretation.
Multivariate logistic regression analysis showed that WBC, RBC, hs-CRP, and NCR were independent predictors of recurrence. LASSO regression selected 11 variables for model development. Among the candidate models, GBM demonstrated the most balanced predictive performance and was therefore selected as the final model. The AUCs of GBM in the training, testing, and validation sets were 0.777, 0.784, and 0.712, respectively. Calibration curves showed acceptable agreement between predicted and observed risks, while decision curve analysis indicated potential clinical benefit within low-to-moderate threshold probability ranges. SHAP analysis identified age, WBC, RBC, NCR, and hs-CRP as the main contributors to model prediction. Restricted cubic spline analysis revealed a significant nonlinear association between NCR and recurrence risk.
WBC, RBC, hs-CRP, and NCR were independently associated with postoperative recurrence of anal fistula. The LASSO-based GBM model demonstrated stable predictive performance and acceptable clinical utility. Routine hematological parameters and inflammation-related indices, particularly NCR, may support individualized recurrence risk stratification and postoperative follow-up.
Yun-Hao Zhou, Da-Wei Wang, Min Tang et al.· Frontiers in Surgery· 0 citations
Related blog posts
MIT News · Artificial Intelligence· news.mit.eduAug 27, 2026
A new machine-learning framework aims to improve the success rate of computational protein design while moving away from results that reproduce sequences found in nature.