Accelerate Literature Icon
Want to do a literature review? Try our new Literature Review workflow

An interpretable machine learning model for predicting the risk of post-endoscopic retrograde cholangiopancreatography complications in elderly patients with choledocholithiasis

  • Abstract
  • Literature Map
  • Similar Papers
Abstract
Translate article icon Translate Article Star icon

Background:Identifying an effective method to predict the risk of post-endoscopic retrograde cholangiopancreatography (ERCP) complications is of significant importance for developing personalized treatment plans and improving outcomes. Previous studies have notable limitations; hence, we have developed and validated an interpretable machine learning model for predicting complications.Methods:We collected data from patients over 60 years old with choledocholithiasis treated at the First Medical Center of Chinese PLA General Hospital between January 2010 and December 2024. We constructed models using nine machine learning methods, including random forest (RF). The predictive performance of the models was compared using evaluation metrics, such as the area under the receiver operating characteristic curve (AUC). The SHapley Additive exPlanations (SHAP) method was used to rank the importance of features and to interpret the final model.Results:Among the nine models, the RF model performed the best. The RF model was able to accurately predict the postoperative complication risk of patients in both the internal validation of the training set [AUC = 0.970 ± 0.021, 95% CI = (0.956, 0.991)] and the test set [AUC = 0.946 ± 0.010, 95% CI = (0.926, 0.996)]. Calibration curves showed a high degree of concordance between predicted and observed risks. The final RF model included nine variables: distal common bile duct angle, cannulation time, INR, number of difficult stone factors, CRP, age, stone count, stent, and length of the distal common bile duct.Conclusion:We have developed and validated an interpretable RF model for predicting the risk of post-ERCP complications in elderly patients with choledocholithiasis, and proposed targeted preventive strategies.

Similar Papers
  • Research Article
  • Cite Count Icon 10
  • 10.1159/000497424
Prediction Model of Cardiac Risk for Dental Extraction in Elderly Patients with Cardiovascular Diseases
  • May 2, 2019
  • Gerontology
  • Min Tang + 5 more

Background: With the rapidly increasing population of elderly people, dental extraction in elderly individuals with cardiovascular diseases (CVDs) has become quite common. The issue of how to assure the safety of elderly patients with CVDs undergoing dental extraction has perplexed dentists and internists for many years. And it is important to derive an appropriate risk prediction tool for this population. Objectives: The aim of this retrospective, observational study was to establish and validate a prediction model based on the random forest (RF) algorithm for the risk of cardiac complications of dental extraction in elderly patients with CVDs. Methods: Between August 2017 and May 2018, a total of 603 patients who fulfilled the inclusion criteria were used to create a training set. An independent test set contained 230 patients between June 2018 and July 2018. Data regarding clinical parameters, laboratory tests, clinical examinations before dental extraction, and 1-week follow-up were retrieved. Predictors were identified by using logistic regression (LR) with penalized LASSO (least absolute shrinkage and selection operator) variable selection. Then, a prediction model was constructed based on the RF algorithm by using a 5-fold cross-validation method. Results: The training set, based on 603 participants, including 282 men and 321 women, had an average participant age of 72.38 ± 8.31 years. Using feature selection methods, 11 predictors for risk of cardiac complications were screened out. When the RF model was constructed, its overall classification accuracy was 0.82 at the optimal cutoff value of 18.5%. In comparison to the LR model, the RF model showed a superior predictive performance. The AUROC (area under the receiver operating characteristic curve) scores of the RF and LR models were 0.83 and 0.80, respectively, in the independent test set. The AUPRC (area under the precision-recall curve) scores of the RF and LR models were 0.56 and 0.35, respectively, in the independent test set. Conclusion: The RF-based prediction model is expected to be applicable for preoperative clinical assessment for preventing cardiac complications in elderly patients with CVDs undergoing dental extraction. The findings may aid physicians and dentists in making more informed recommendations to prevent cardiac complications in this patient population.

  • Research Article
  • Cite Count Icon 5
  • 10.1080/00015385.2025.2481662
Application of an interpretable machine learning method to predict the risk of death during hospitalization in patients with acute myocardial infarction combined with diabetes mellitus
  • Apr 7, 2025
  • Acta Cardiologica
  • Zhijun Bu + 12 more

Background Predicting the prognosis of patients with acute myocardial infarction (AMI) combined with diabetes mellitus (DM) is crucial due to high in-hospital mortality rates. This study aims to develop and validate a mortality risk prediction model for these patients by interpretable machine learning (ML) methods. Methods Data were sourced from the Medical Information Mart for Intensive Care IV (MIMIC-IV, version 2.2). Predictors were selected by Least absolute shrinkage and selection operator (LASSO) regression and checked for multicollinearity with Spearman’s correlation. Patients were randomly assigned to training and validation sets in an 8:2 ratio. Seven ML algorithms were used to construct models in the training set. Model performance was evaluated in the validation set using metrics such as area under the curve (AUC) with 95% confidence interval (CI), calibration curves, precision, recall, F1 score, accuracy, negative predictive value (NPV), and positive predictive value (PPV). The significance of differences in predictive performance among models was assessed utilising the permutation test, and 10-fold cross-validation further validated the model’s performance. SHapley Additive exPlanations (SHAP) and Local Interpretable Model-agnostic Explanations (LIME) were applied to interpret the models. Results The study included 2,828 patients with AMI combined with DM. Nineteen predictors were identified through LASSO regression and Spearman’s correlation. The Random Forest (RF) model was demonstrated the best performance, with an AUC of 0.823 (95% CI: 0.774–0.872), high precision (0.867), accuracy (0.873), and PPV (0.867). The RF model showed significant differences (p < 0.05) compared to the K-Nearest Neighbours and Decision Tree models. Calibration curves indicated that the RF model’s predicted risk aligned well with actual outcomes. 10-fold cross-validation confirmed the superior performance of RF model, with an average AUC of 0.828 (95% CI: 0.800–0.842). Significant Variables in RF model indicated that the top eight significant predictors were urine output, maximum anion gap, maximum urea nitrogen, age, minimum pH, maximum international normalised ratio (INR), mean respiratory rate, and mean systolic blood pressure. Conclusion This study demonstrates the potential of ML methods, particularly the RF model, in predicting in-hospital mortality risk for AMI patients with DM. The SHAP and LIME methods enhance the interpretability of ML models.

  • Research Article
  • Cite Count Icon 1
  • 10.1200/po-25-00192
Interpretable Machine Learning Models for Predicting Lateral Pelvic Lymph Node Metastasis in Rectal Cancer: A Chinese Multicenter Retrospective Study.
  • Sep 1, 2025
  • JCO precision oncology
  • Tixian Xiao + 9 more

Internal iliac and obturator lymph nodes are common sites of metastasis in rectal cancer. This study developed a machine learning (ML) model using clinical data to predict lymph node metastasis and applied the Shapley Additive explanations (SHAP) method for interpretation. Retrospectively, data from patients with rectal cancer at four Chinese centers-who underwent total mesorectal excision and lateral pelvic lymph node dissection without neoadjuvant therapy-were collected. Two centers provided training/test sets (3:1 ratio) and two centers supplied external validation. Lymph node enlargement was determined by imaging and confirmed by pathology. Five ML models were evaluated by AUC, accuracy, and F1 score. Key features included demographics, tumor stage, tumor-to-anal verge distance, imaging measurements, tumor histological differentiation, preoperative carcinoembryonic antigen, and carbohydrate antigen 19-9. SHAP was used to assess feature importance. Of the 411 cases (174 positives) in the training/test sets and 109 cases (43 positives) in external validation, the random forest (RF) model ranked second in terms of AUC and accuracy in the training set (0.999, 0.995), whereas it achieved the highest AUC and accuracy (0.877 and 0.788) in the test set. In the external validation, the RF model outperformed all other ML models (AUC of 0.899, accuracy of 0.827). Overall, the RF model demonstrates the superior overall performance. According to the SHAP analysis, the most important predictors of internal iliac and obturator lymph node metastasis were, in descending order, the short-axis diameter of enlarged lymph nodes, regional lymph node metastasis, and tumor-to-anal verge distance. At the individual patient level, SHAP force plots provided explanations of the RF model predictions for internal iliac and obturator lymph node metastasis. An interpretable ML model was developed that accurately predicts internal iliac and obturator lymph node metastasis using clinical data. SHAP analysis enhances understanding of feature contributions, supporting personalized treatment planning.

  • Research Article
  • Cite Count Icon 4
  • 10.3389/fphys.2025.1542240
Web-based machine learning application for interpretable prediction of prolonged length of stay after lumbar spinal stenosis surgery: a retrospective cohort study with explainable AI
  • Feb 19, 2025
  • Frontiers in Physiology
  • Paierhati Yasheng + 5 more

ObjectivesLumbar spinal stenosis (LSS) is an increasingly important issue related to back pain in elderly patients, resulting in significant socioeconomic burdens. Postoperative complications and socioeconomic effects are evaluated using the clinical parameter of hospital length of stay (LOS). This study aimed to develop a machine learning-based tool that can calculate the risk of prolonged length of stay (PLOS) after surgery and interpret the results.MethodsPatients were registered from the spine surgery department in our hospital. Hospital stays greater than or equal to the 75th percentile for LOS was considered extended PLOS after spine surgery. We screened the variables using the least absolute shrinkage and selection operator (LASSO) and permutation importance value and selected nine features. We then performed hyperparameter selection via grid search with nested cross-validation. Receiver operating characteristics curve, calibration curve and decision curve analysis was carried out to assess model performance. The result of the final selected model was interpreted using Shapley Additive exPlanations (SHAP), and Local Interpretable Model-agnostic Explanations (LIME) were used for model interpretation. To facilitate model utilization, a web application was deployed.ResultsA total of 540 patients were involved, and several features were finally selected. The final optimal random forest (RF) model achieved an area under the curve (ROC) of 0.93 on the training set and 0.83 on the test set. Based on both SHAP and LIME analyses, intraoperative blood loss emerged as the most significant contributor to the outcome.ConclusionMachine learning in association with SHAP and LIME can provide a clear explanation of personalized risk prediction, and spine surgeons can gain a perceptual grasp of the impact of important model components. Utilization and future clinical research of our RF model are made simple and accessible through the web application.

  • PDF Download Icon
  • Research Article
  • Cite Count Icon 52
  • 10.1038/s41398-024-02762-w
Prediction models for postoperative delirium in elderly patients with machine-learning algorithms and SHapley Additive exPlanations
  • Jan 25, 2024
  • Translational psychiatry
  • Yuxiang Song + 10 more

Postoperative delirium (POD) is a common and severe complication in elderly patients with hip fractures. Identifying high-risk patients with POD can help improve the outcome of patients with hip fractures. We conducted a retrospective study on elderly patients (≥65 years of age) who underwent orthopedic surgery with hip fracture between January 2014 and August 2019. Conventional logistic regression and five machine-learning algorithms were used to construct prediction models of POD. A nomogram for POD prediction was built with the logistic regression method. The area under the receiver operating characteristic curve (AUC-ROC), accuracy, sensitivity, and precision were calculated to evaluate different models. Feature importance of individuals was interpreted using Shapley Additive Explanations (SHAP). About 797 patients were enrolled in the study, with the incidence of POD at 9.28% (74/797). The age, renal insufficiency, chronic obstructive pulmonary disease (COPD), use of antipsychotics, lactate dehydrogenase (LDH), and C-reactive protein are used to build a nomogram for POD with an AUC of 0.71. The AUCs of five machine-learning models are 0.81 (Random Forest), 0.80 (GBM), 0.68 (AdaBoost), 0.77 (XGBoost), and 0.70 (SVM). The sensitivities of the six models range from 68.8% (logistic regression and SVM) to 91.9% (Random Forest). The precisions of the six machine-learning models range from 18.3% (logistic regression) to 67.8% (SVM). Six prediction models of POD in patients with hip fractures were constructed using logistic regression and five machine-learning algorithms. The application of machine-learning algorithms could provide convenient POD risk stratification to benefit elderly hip fracture patients.

  • Research Article
  • Cite Count Icon 1
  • 10.3389/fmed.2025.1638097
Machine learning-based predictive model for acute pancreatitis-associated lung injury: a retrospective analysis
  • Aug 12, 2025
  • Frontiers in Medicine
  • Zhaohui Du + 9 more

BackgroundAcute Pancreatitis-Associated Lung Injury (APALI) is one of the most severe and life-threatening systemic complications in acute pancreatitis patients, with high rates of morbidity and mortality. This study aims to develop a prediction model for the diagnosis of APALI based on machine learning algorithms.MethodsThis study included data from the First Affiliated Hospital of Bengbu Medical College (July 2012 to June 2022), which were randomly categorized into the training and testing set. And data from the Second Affiliated Hospital of Zhejiang University (January 2018 to April 2023) served as the external validation set. LASSO regression was applied to eliminate irrelevant or highly collinear independent variables. Six machine learning models were constructed, with evaluation metrics including Area Under Curve (AUC), accuracy, sensitivity, specificity, F1 score, and recall. The impact of model features was analyzed using SHapley Additive exPlanations (SHAP).ResultsA total of 1,975 patients with acute pancreatitis were randomly assigned to a training set (1,480 patients) and a testing set (495 patients). In the training set, 480 cases (32.43%) were diagnosed with APALI. The eXtreme Gradient Boosting (XGBoost) and Random Forest (RF) models demonstrated the best predictive performance, achieving the highest AUC (0.92 and 0.914, respectively), along with higher accuracy, F1 score, and recall in the testing set. Six particularly influential factors were identified and ranked as follows: CRP, BMI, neutrophil, calcium, lactate, and neutrophil-to-albumin ratio (NAR). The global interpretability of the XGBoost and RF models, along with these six features, is shown in the SHAP summary plot. These two models were selected as the optimal models for the development of an online calculator for clinical applications and risk stratification.ConclusionWe developed and internally validated a machine learning model to predict APALI, showing strong performance in our study population. To support further research and clinical use, we created an open-access web-based risk calculator. Prospective multicenter validation is needed to confirm generalizability. If successful, the tool may support early risk identification and guide interventions to prevent APALI.

  • Research Article
  • Cite Count Icon 35
  • 10.1097/tp.0000000000002923
Seeing the Forest for the Trees: Random Forest Models for Predicting Survival in Kidney Transplant Recipients.
  • May 1, 2020
  • Transplantation
  • Ruth Sapir-Pichhadze + 1 more

Risk prediction plays an important role in clinical transplantation research. Traditionally, most risk models have been based on regression models.1 Although useful to help understand relationships between predictors and outcomes, these statistical methods can typically evaluate only a small number of predictors, which are assumed to affect everyone in the same way, and uniformly throughout the participants' lifespan. These methods have several limitations,2 including the inability to analyze nonlinear relationships, the requirement of setting a level of binary significance, impracticality for analyzing large datasets, and vulnerability to bias secondary to variable selection and/or omission of relevant confounders. With the emergence of P4 (Predictive, Preventive, Personalized, and Participatory) and Precision Medicine, artificial intelligence and machine learning methods have come to attention as methods aimed at solving the challenges in analysis not well addressed by regression approaches. Machine learning methods provide algorithms to understand patterns from large, complex, and heterogeneous data.3 Of the machine learning methods, recursive partitioning, and especially random forests, can deal with large numbers of predictor variables even in the presence of complex interactions.2,4 These methods have been applied successfully in genetics, clinical research, and bioinformatics. In this issue of Transplantation, Scheffner et al report on the development and internal validation of a random forest prediction model for patient survival.5 Random forest models are composed of a collection of decision trees. In the process of building each decision tree, different random subsets of the variables from the training dataset are selected to establish how best to partition the dataset at each node.6 Random forest models are considered less vulnerable to overfitting the training dataset given the large number of trees built, making each tree an independent model. The lower likelihood of bias is a result of bootstrapping several trees over randomly selected subsets of variables and subsamples of data.6 Random forest models require little preprocessing of data; the data need not be normalized; and the approach is resilient to outliers. While missing data will be a challenge when trying to draw clinical inferences from standard statistical models, machine learning methods tend to make fewer assumptions about the underlying data and, thus, are less vulnerable to the challenges associated with violation of those assumptions. Relying on fewer assumptions than regression analysis, machine learning methods have been shown to deliver more robust predictions. Scheffner and colleagues5 split a retrospective cohort of kidney transplant recipients with posttransplantation protocol biopsies into training and validation datasets (Figure 2A and B). Using all pretransplant and 3- and 12-months posttransplant variables, the obtained models showed good performance to predict death (concordance index: 0.77–0.78). Validation showed a concordance index of 0.76 and good discrimination of risks by the models, despite substantial differences in clinical variables and the derivation dataset representing an earlier era (2000–2007) than the validation dataset (2008–2013). To contrast with outputs of multivariable regression models using the same datasets, see Tables 2 and 3 and nomograms predicting mortality risk using estimators from multivariable Cox models (Figure 3) in Abeling et al.7 Random survival forests also inform on the importance of descriptive variables.6 Scheffner found the potentially modifiable (and highly correlated) graft rejection treatment and urinary tract infection to be important predictors of patient survival in addition to established factors like age, cardiovascular disease, diabetes, and graft function (Figure 3A and B).5 Many of the predictors retained in multivariable regression models7 were also deemed important in random forest survival analyses.5 To validate selected predictors and model construction, it is important to pursue external validation with independent datasets. Random survival forests may complement regression analyses when handling highly correlated complex survival data. Opportunities for application (and limitations) of each of the regression and random survival forests for prediction are summarized in Table 1.TABLE 1.: Regression and random survival forests for survival analysisPredictive models in transplantation and donation help risk stratify patients and could improve quality of healthcare delivery as well as patient outcomes. The increasing interest in these tools warrants a better understanding of their challenges and limitations.8 First, highly predictive variables may not necessarily be causally related to the outcomes of interest. Second, the success of machine learning models depends on the relationship between predictors and outcome being represented in training/validation datasets, the number of observations and features, selection and parameterization of features, and the algorithm chosen for the model. Careful variable definition (eg, urinary tract infection) is necessary. Presence of highly correlated linear and nonlinear relationships between independent variables may warrant mechanisms for removal of the correlated variables. Model performance may also be compromised when studying rare outcomes.4 Inevitably, generalizability of machine learning models may be limited when the clinical context, local factors (including patient/physician preferences, health systems, and care standards), and therapeutic strategies vary. To enable assessment of model validity, correct interpretation of model outputs, replication, and future knowledge synthesis, it is vital that the transplantation and donation community promote adherence to guidelines on the dissemination and reporting of machine learning models.8,9 Authors should be encouraged to report all model parameters, transformations applied to raw data, sampling methods, and random number generator seeds. Whenever possible, algorithms and associated code should be released in public software archive domains. There is a need for new models of health data ownership with rights to the individual, highly secure data repositories, government legislation for data sharing, and usage policies to ensure privacy and data security. Moreover, with wide uptake of machine learning and artificial intelligence tools, the scale of iatrogenic risks and liabilities related to their application, in contrast to the implications of a single doctor's mistake for a given patient, also warrant assessment.10 Most practice guidelines are geared toward the "average patient." Machine learning tools can capture the complexity of individual patients' characteristics and aid transplant clinicians with patient-specific care decisions. As these tools become more prevalent, it is important to develop best practice guidelines and ensure there is regulatory oversight on their development and application.

  • Research Article
  • 10.3389/fonc.2025.1661695
Interpretable machine learning models based on multi-dimensional fusion data for predicting positive surgical margins in robot-assisted radical prostatectomy: a retrospective study
  • Oct 3, 2025
  • Frontiers in Oncology
  • Zhangcheng Liu + 16 more

ObjectiveThis study aimed to develop and validate interpretable machine learning (ML) models based on multi-dimensional fusion data for predicting positive surgical margins (PSM) in robot-assisted radical prostatectomy (RARP).MethodsPatients who underwent RARP at our institution between January 2016 and July 2025 were enrolled. Demographic, clinical, biopsy pathology data, and MRI-derived anatomical features (measured using ITK-SNAP on axial, sagittal, and coronal planes) were collected. Feature selection was performed using intraobserver and interobserver correlation coefficients (ICCs), low-variance filtering, univariable logistic regression, Spearman’s correlation analysis, the least absolute shrinkage and selection operator (LASSO) algorithm, and the Boruta algorithm. Six ML models were constructed, with performance evaluated using area under the curve (AUC), calibration curves, and decision curve analyses (DCA) to identify the optimal model. Five-fold and ten-fold cross-validation were used to assess the optimal model’s generalizability, and its interpretability was evaluated via Shapley Additive exPlanations (SHAP) analysis.ResultsA total of 347 patients were included, comprising a training set (n=193, January 2016–December 2024), validation set (n=84, January 2016–December 2024), and test set (n=70, January 2025–July 2025). From 164 initial features, 7 key features were retained through a four-step screening. The Random Forest (RF) model outperformed other models, achieving AUCs of 0.99 (95% CI: 0.97–1.00) in the training set, 0.88 (95% CI: 0.80–0.95) in the validation set, and 0.97 (95% CI: 0.94–1.00) in the test set. Calibration curve and decision curve analyses confirmed its strong clinical utility. Five-fold cross-validation for the RF model showed fold-specific AUCs of 0.82–0.92, with a mean AUC of 0.87 (95% CI: 0.84–0.90). Ten-fold cross-validation showed fold-specific AUCs of 0.80–0.99, with a mean AUC of 0.88 (95% CI: 0.83–0.93). SHAP analysis revealed five novel spatial anatomical features (such as Sagittal plane-posterior spatial anatomical structure index, Coronal plane-Left anatomical structure interval) were negatively associated with PSM risk, while the number of positive biopsy cores and clinical tumor stage were positively associations.ConclusionsMulti-dimensional fusion data combined with ML models improves PSM prediction accuracy in RARP. The RF model, with excellent performance and interpretability, shows promise for preoperative PSM risk stratification, facilitates optimized clinical decision-making, and supports personalized treatment discussions during preoperative planning, but requires prospective and external validation before clinical implementation.

  • Research Article
  • Cite Count Icon 2
  • 10.4103/sjg.sjg_231_25
A risk score model for post-endoscopic retrograde cholangiopancreatography complications in elderly patients with choledocholithiasis
  • Aug 28, 2025
  • Saudi Journal of Gastroenterology : Official Journal of the Saudi Gastroenterology Association
  • Guanjun Zhang + 8 more

Background:Endoscopic retrograde cholangiopancreatography (ERCP) is the preferred treatment for choledocholithiasis. This study aimed to identify the independent risk factors for post-ERCP complications in elderly patients with choledocholithiasis and establish a risk score model.Methods:This study enrolled patients over the age of 75 with choledocholithiasis, who underwent ERCP at two hospitals. Multivariate logistic regression analysis was used to identify predictive factors and establish a weighted risk score model, which was then externally validated.Results:Five factors (Charlson comorbidity index ≥3, aspartate aminotransferase >the upper limit of normal, endoscopic papillary large balloon dilation, cannulation time >10 min, and multiple large stones) were identified as risk factors for post-ERCP complications. A seven-point score model was established, with a score of three or above considered high risk and less than three considered low risk. The areas under the receiver operating characteristic curve for the derivation and validation cohorts were 0.791 (95% CI, 0.742–0.794) and 0.876 (95% CI, 0.793–0.879), respectively, with sensitivities of 0.755 (95% CI, 0.660–0.831) and 0.901 (95% CI, 0.698–0.972), and specificities of 0.826 (95% CI, 0.791–0.856) and 0.854 (95% CI, 0.761–0.914), respectively.Conclusions:An easy-to-use score model was successfully derived to help predict the risk of post-ERCP complications in elderly patients with choledocholithiasis.

  • PDF Download Icon
  • Research Article
  • Cite Count Icon 12
  • 10.1186/s12911-023-02411-0
Prediction of postoperative infectious complications in elderly patients with colorectal cancer: a study based on improved machine learning
  • Jan 6, 2024
  • BMC Medical Informatics and Decision Making
  • Yuan Tian + 5 more

BackgroundInfectious complications after colorectal cancer (CRC) surgery increase perioperative mortality and are significantly associated with poor prognosis. We aimed to develop a model for predicting infectious complications after colorectal cancer surgery in elderly patients based on improved machine learning (ML) using inflammatory and nutritional indicators.MethodsThe data of 512 elderly patients with colorectal cancer in the Third Affiliated Hospital of Anhui Medical University from March 2018 to April 2022 were retrospectively collected and randomly divided into a training set and validation set. The optimal cutoff values of NLR (3.80), PLR (238.50), PNI (48.48), LCR (0.52), and LMR (2.46) were determined by receiver operating characteristic (ROC) curve; Six conventional machine learning models were constructed using patient data in the training set: Linear Regression, Random Forest, Support Vector Machine (SVM), BP Neural Network (BP), Light Gradient Boosting Machine (LGBM), Extreme Gradient Boosting (XGBoost) and an improved moderately greedy XGBoost (MGA-XGBoost) model. The performance of the seven models was evaluated by area under the receiver operator characteristic curve, accuracy (ACC), precision, recall, and F1-score of the validation set.ResultsFive hundred twelve cases were included in this study; 125 cases (24%) had postoperative infectious complications. Postoperative infectious complications were notably associated with 10 items features: American Society of Anesthesiologists scores (ASA), operation time, diabetes, presence of stomy, tumor location, NLR, PLR, PNI, LCR, and LMR. MGA-XGBoost reached the highest AUC (0.862) on the validation set, which was the best model for predicting postoperative infectious complications in elderly patients with colorectal cancer. Among the importance of the internal characteristics of the model, LCR accounted for the highest proportion. Conclusions: This study demonstrates for the first time that the MGA-XGBoost model with 10 risk factors might predict postoperative infectious complications in elderly CRC patients.

  • Research Article
  • 10.1371/journal.pone.0331857
Optimizing clinical prediction model for new-onset atrial fibrillation in critically ill patient: Based on machine learning
  • Sep 11, 2025
  • PLOS One
  • Da-Cheng Wang + 3 more

BackgroundNew-onset atrial fibrillation (NOAF) increases the risk of embolism and sudden death in critically ill patients; however, limited data exist attempting to identify modifiable risk factors and predict the incidence of NOAF. We aimed to investigate the risk factors for NOAF and develop an optimized clinical prediction model based on machine learning algorithms.Materials and methodsData from patients admitted to the intensive care unit (ICU) of the Affiliated Hospital of Nanjing University of Chinese Medicine from August 2019 to January 2022 were retrospectively analyzed. LASSO regression and Random Forest (RF) algorithms were used to screen predictive variables. Logistic Regression, RF, Gradient Boosting and Support Vector Machine models were constructed to evaluate the recognition ability of different machine learning algorithms. The confusion matrix and calibration curve were used to assess the degree of accuracy of the four models. Decision curve analysis (DCA) was conducted to evaluate the utility of the model in decision-making. The net reclassification index (NRI) and integrated discrimination improvement (IDI) were also calculated to evaluate the performance of the models. The learning curves of the four models were plotted to evaluate the precision of different models. The SHapley Additive exPlanations (SHAP) was used to explain the supreme-performing model.ResultsIn total, 417 patients were enrolled in the study, and 333 patients were allocated to the training group and 84 to the validation group. The baseline characteristic distributions were similar between the two groups. Age, heart rate, mean arterial pressure, activated partial thromboplastin time, and brain natriuretic peptide were revealed as independent predictors of NOAF by LASSO regression and the RF algorithm. The RF model had the best performance, with the area under the receiver operator characteristic curve (AUROC) of 0.758, the area under the precision-recall curve (AUPRC) of 0.524, and accuracy of 0.735 in the training set, paralleled by AUROC of 0.796, AUPRC of 0.686, and accuracy of 0.702 in the validation set. The confusion matrix and calibration curves showed that RF had the best performance. DCAs also showed that the RF model provided the highest net benefit in the clinical setting. The NRI results showed that the RF significantly improved reclassification ability compared to the baseline model (NRI = 0.38). The IDI results further demonstrated a moderate improvement in discrimination ability for the RF (IDI = 0.033) compared to the baseline. The learning curves revealed that RF also showed superior performance. SHAP could be used visualized individual NOAF risk predicted by the model.ConclusionsThe RF model exhibited the best performance in predicting NOAF in critically ill patients and has the potential to help clinicians identify high-risk patients and guide clinical decision making.

  • Research Article
  • 10.1097/md.0000000000048408
A machine learning model incorporating the globulin-to-platelet index for predicting severe fibrosis in autoimmune hepatitis: A retrospective and prospective validation study
  • May 8, 2026
  • Medicine
  • Haiping Zhang + 10 more

Accurate staging of liver fibrosis in autoimmune hepatitis (AIH) remains challenging due to the invasive nature and sampling limitations of liver biopsy. This study aimed to identify readily available predictors of severe fibrosis and to develop an AIH-specific noninvasive machine-learning model. This two-stage study retrospectively enrolled 208 patients with biopsy-confirmed AIH, with prospective validation in 26 additional patients. Transient elastography (TE) was performed in 110 retrospective and 12 prospective patients. Severe fibrosis was defined as Scheuer stages S3 to S4. Candidate variables underwent univariable and multivariable logistic regression with collinearity control. A random forest (RF) model was trained on the independent predictors and evaluated by the area under the receiver operating characteristic curve (AUROC), calibration, and decision curve analysis. Shapley Additive exPlanations were used for interpretability. Inflammatory activity was graded by Scheuer and prespecified for subgroup analyses (G0–G2 vs G3–G4). A TE-inclusive RF model was also developed in the TE subgroup. The globulin-to-platelet index, international normalized ratio, and blood urea nitrogen were identified as independent predictors of severe fibrosis in AIH. The RF model based on these variables yielded AUROCs of 0.863 (95% confidence interval [CI], 0.802–0.917) in the training set, 0.747 (95% CI, 0.602–0.863) in the test set, and 0.784 (95% CI, 0.556–0.959) in the prospective cohort. Stratified by inflammatory grade, AUROCs were 0.842 (95% CI, 0.757–0.914) in G0–G2 and 0.814 (95% CI, 0.726–0.886) in G3–G4. In contrast, the aspartate aminotransferase-to-platelet ratio index and fibrosis-4 index performed poorly overall and deteriorated further under moderate-to-severe inflammation. In the TE subgroup, the RF model outperformed TE alone (AUROC, 0.786 vs 0.682), and performance improved further when TE was integrated (AUROC, 0.898 [95% CI, 0.841–0.949]). Globulin-to-platelet index, international normalized ratio, and blood urea nitrogen were independent predictors of severe fibrosis in AIH. An RF model constructed from these markers provided a robust, noninvasive tool whose performance was preserved across inflammatory grades and was further enhanced by incorporating TE.

  • Research Article
  • 10.1182/blood-2025-5784
Identification of dynamic markers for multiple myeloma in patients with MGUS and T2DM
  • Nov 3, 2025
  • Blood
  • Mei Wang + 3 more

Identification of dynamic markers for multiple myeloma in patients with MGUS and T2DM

  • Research Article
  • Cite Count Icon 3
  • 10.13168/agg.2024.0019
GIS-based landslide susceptibility assessment using random forest and support vector machine models: A case study for Chin State, Myanmar
  • Sep 24, 2024
  • Acta Geodynamica et Geomaterialia
  • Soe Hlaing Tun + 2 more

Chin State in Myanmar experiences frequent landslides annually. This research aimed to construct GIS-based landslide susceptibility maps (LSMs) with two kinds of machine learning models, namely random forest (RF) and support vector machine (SVM). Firstly, a landslide inventory map was constructed by containing 213 landslide locations and randomly chosen 213 non-landslide locations; these location points were randomly divided into the training set (70 %) for the landslide susceptibility prediction model and the testing set (30 %) for the model validation. Secondly, twenty-one landslide conditioning factors were selected, and frequency ratio analysis was used to evaluate the relationship between each class of factors and landslide occurrences. Then, landslide susceptibility prediction modeling by RF and SVM models. Finally, the performance of the two models was evaluated with performance metrics (precision, recall, F1-Score, and accuracy), receiver operating characteristic (ROC) curves, and area under the ROC curve (AUC values). The RF model demonstrated superior performance across performance evaluation metrics, with a precision of 0.864, recall of 0.919, F1-Score of 0.891, and an accuracy of 0.894 on the training set, compared to the SVM model's precision of 0.854, recall of 0.807, F1-Score of 0.830, and accuracy of 0.825. The model validation by the testing set further confirmed that the RF model showed a precision of 0.839, recall of 0.897, F1-Score of 0.867, and an accuracy of 0.871, while the SVM model had a precision of 0.839, recall of 0.839, F1-Score of 0.839, and an accuracy of 0.839. Also, the results of AUC values showed that the RF model (training set AUC = 0.94, testing set AUC = 0.92), and SVM model (training set AUC = 0.89, testing set AUC = 0.88), respectively. Hence, these two landslide susceptibility prediction models demonstrated satisfactory results and good accuracy for LSMs in this research area, and the LSM from the RF model is better than the SVM model according to performance metrics and AUC values results. The resulting maps provide useful information on the likelihood of landslide occurrence, facilitating decision-making in land use planning and disaster management.

  • Research Article
  • Cite Count Icon 39
  • 10.1016/j.enconman.2024.118808
Machine learning development to predict the electrical efficiency of photovoltaic-thermal (PVT) collector systems
  • Jul 23, 2024
  • Energy Conversion and Management
  • Hossein Gharaee + 2 more

According to the increasing rate of renewable energy use, especially solar energy, photovoltaic-thermal (PVT) solar panels are getting more attention and they have become significantly important for extracting heat and electricity from solar energy. In this regard, the theoretical equations used to calculate the PVT performance lack accuracy, and many uncertainties and calculation impediments result in deviation from the experimental outputs. Therefore, machine learning (ML) methods can overcome these disadvantages and can be used to predict PVT efficiency more reliably and effectively. In this study, three ML methods, namely, multilayer perceptron (MLP), random forest (RF), and support vector regression (SVR) were implemented to build trained models predicting the electrical efficiency of PVTs. These models are based on more than 380 datasets that were extracted from the literature where all PVT benches use water as their working fluid. The models considered mass flow rate, solar radiation, ambient temperature, wind speed, fluid inlet temperature, PVT surface area, and pipe inner diameter as input variables. The optimized models were attained through rigorous hyperparameters variations through random, coarse and fine grid searches. The results showed that the RF model has a root mean squared error (RMSE) of 0.2233, 0.4638, and 0.671 for training, testing, and validation datasets respectively, and training coefficient of determination (R-square), with a value of 0.9862, providing relatively high prediction accuracy for the electrical efficiency. Next, the MLP model with a testing R-square value of 0.8775 and the SVR model with a testing R-square value of 0.7639 can predict electrical efficiencies. Further, the output results are visualized using explainable artificial intelligence (AI) method such as SHapley Additive exPlanations (SHAP) which declared that mass flow rate, pipe inner diameter, and wind speed are the most effective variables in the RF model. For validation purposes, 46 datasets characterized with at least one out-domain parameter are used to test the performance of our models. For out-of-range input variables, the RF model most accurately predicted the electrical efficiency for cell absorptance and inlet fluid temperature variables, outperforming the MLP and SVR models. Meanwhile, the SVR model showed superior accuracy over both RF and MLP when dealing with datasets featuring solar radiation values beyond the training domain.

Save Icon
Up Arrow
Open/Close
Notes

Save Important notes in documents

Highlight text to save as a note, or write notes directly

You can also access these Documents in Paperpal, our AI writing tool

Powered by our AI Writing Assistant