Accelerate Literature Icon
Want to do a literature review? Try our new Literature Review workflow

Machine learning model integrating CT radiomics and circulating microRNAs to predict residual disease histology in metastatic non-seminoma testicular cancer (mNSTC).

  • Abstract
  • Literature Map
  • Similar Papers
Abstract
Translate article icon Translate Article Star icon

647 Background: The primary treatment of most mNSTC is chemotherapy followed by surgery if the residual disease (RD) is >1 cm. However, conventional imaging lacks the specificity to characterize the tissue, often leading to overtreatment. This study hypothesizes that integrating CT-driven radiomics features with plasma miR371 and miR375 will enhance the predictive accuracy of Machine Learning (ML) models to predict teratoma, viable germ cell (vGCT) and fibrosis/necrosis (F/N) in mNSTC patients with RD. Methods: 111 lesions from52 patients, including residual teratoma (n=57), F/N (n=33), vGCT (n=10), and additional seminoma (n=11) for training purposes were included, split into training (N=78) and test cohorts (N=33). Lesions were lymph nodes (n=87), lung (n=21), and brain (n=3) with a median size of 1.6 cm (Q1-Q3 interval=1.2-2.73 cm). 3D Slicer version 5.6.1 was used to segment the RD > 1 cm (short axis) and extract radiomics features. Plasma miRNA levels before resection were measured by RT-PCR. Random Forest (RF), Support Vector Machine (SVM), Gradient Boosting (GB), and CatBoost (CB) ML models were evaluated to define the operating characteristics of radiomics alone (R-only) and in combination with miR371 (371) and/or miR375 (375) levels in predicting teratoma, vGCT and F/N. Results: For predicting teratoma, the best models were RF (R+375 and R+371+375), CB (R+371+375), and GB (R+371 and R+371+375). While adding miR371 or miR375 to R-only slightly improved AUC across models, the best results were achieved with the R+375+371 dataset. CB achieved AUCs ranging from 0.94 to 0.97 in training and 0.81 to 0.93 in test sets, with its highest AUC of 0.93 (95% CI: 0.78-0.97) on the R+375+371 dataset to differentiate all three classes. Similarly, GB demonstrated strong performance, achieving its highest AUC of 0.93 (95% CI: 0.79-0.96) on the R+375+371 dataset (Table). Conclusions: Integration of plasma miR371, miR375 and radiomics improved accuracy of predicting histologies across all ML models. These methods could be used to characterize the histology of RD in mNSTC patients to better inform treatment decisions. Further refinement, including incorporation of histological findings of the primary tumor, will be reported. AUC values of different ML algorithms on training and test sets. TRAINING SET TEST SET Model ±SD R R+375 R+371 R+375+371 Model (95% CI) R R+375 R+371 R+375+371 RF 0.93±0.05 0.95±0.04 0.95±0.03 0.96±0.04 RF 0.8(0.59-0.89) 0.85(0.72-0.93) 0.87(0.76-0.95) 0.91(0.78-0.95) SVM 0.84±0.06 0.84±0.09 0.89±0.11 0.89±0.09 SVM 0.72(0.54-0.80) 0.74(0.56-0.82) 0.83(0.69-0.92) 0.84(0.76-0.94) GB 0.94±0.04 0.91±0.08 0.95±0.05 0.97±0.03 GB 0.84(0.61-0.96) 0.89(0.77-0.97) 0.89(0.79-0.96) 0.93(0.79-0.96) CB 0.95±0.03 0.94±0.03 0.94±0.04 0.97±0.03 CB 0.81(0.6-0.93) 0.86(0.73-0.94) 0.89(0.78-0.97) 0.93(0.78-0.97)

Similar Papers
  • Research Article
  • Cite Count Icon 9
  • 10.1007/s11307-023-01823-8
Application of Machine Learning Analyses Using Clinical and [18F]-FDG-PET/CT Radiomic Characteristics to Predict Recurrence in Patients with Breast Cancer.
  • May 16, 2023
  • Molecular imaging and biology
  • Kodai Kawaji + 8 more

To develop and identify machine learning (ML) models using pretreatment clinical and 2-deoxy-2-[18F]fluoro-D-glucose positron emission tomography ([18F]-FDG-PET)-based radiomic characteristics to predict disease recurrences in patients with breast cancers who underwent surgery. This retrospective study included 112 patients with 118 breast cancer lesions who underwent [18F]-FDG-PET/ X-ray computed tomography (CT) preoperatively, and these lesions were assigned to training (n=95) and testing (n=23) cohorts. A total of 12 clinical and 40 [18F]-FDG-PET-based radiomic characteristics were used to predict recurrences using 7 different ML algorithms, namely, decision tree, random forest (RF), neural network, k-nearest neighbors, naive Bayes, logistic regression, and support vector machine (SVM) with a 10-fold cross-validation and synthetic minority over-sampling technique. Three different ML models were created using clinical characteristics (clinical ML models), radiomic characteristics (radiomic ML models), and both clinical and radiomic characteristics (combined ML models). Each ML model was constructed using the top ten characteristics ranked by the decrease in Gini impurity. The areas under ROC curves (AUCs) and accuracies were used to compare predictive performances. In training cohorts, all 7 ML algorithms except for logistic regression algorithm in the radiomics ML model (AUC = 0.760) achieved AUC values of >0.80 for predicting recurrences with clinical (range, 0.892-0.999), radiomic (range, 0.809-0.984), and combined (range, 0.897-0.999) ML models. In testing cohorts, the RF algorithm of combined ML model achieved the highest AUC and accuracy (95.7% (22/23)) with similar classification performance between training and testing cohorts (AUC: training cohort, 0.999; testing cohort, 0.992). The important characteristics for modeling process of this RF algorithm were radiomic GLZLM_ZLNU and AJCC stage. ML analyses using both clinical and [18F]-FDG-PET-based radiomic characteristics may be useful for predicting recurrence in patients with breast cancers who underwent surgery.

  • PDF Download Icon
  • Research Article
  • Cite Count Icon 40
  • 10.3390/ijgi9040276
Comparing Machine Learning Models and Hybrid Geostatistical Methods Using Environmental and Soil Covariates for Soil pH Prediction
  • Apr 23, 2020
  • ISPRS International Journal of Geo-Information
  • Panagiotis Tziachris + 4 more

In the current paper we assess different machine learning (ML) models and hybrid geostatistical methods in the prediction of soil pH using digital elevation model derivates (environmental covariates) and co-located soil parameters (soil covariates). The study was located in the area of Grevena, Greece, where 266 disturbed soil samples were collected from randomly selected locations and analyzed in the laboratory of the Soil and Water Resources Institute. The different models that were assessed were random forests (RF), random forests kriging (RFK), gradient boosting (GB), gradient boosting kriging (GBK), neural networks (NN), and neural networks kriging (NNK) and finally, multiple linear regression (MLR), ordinary kriging (OK), and regression kriging (RK) that although they are not ML models, they were used for comparison reasons. Both the GB and RF models presented the best results in the study, with NN a close second. The introduction of OK to the ML models’ residuals did not have a major impact. Classical geostatistical or hybrid geostatistical methods without ML (OK, MLR, and RK) exhibited worse prediction accuracy compared to the models that included ML. Furthermore, different implementations (methods and packages) of the same ML models were also assessed. Regarding RF and GB, the different implementations that were applied (ranger-ranger, randomForest-rf, xgboost-xgbTree, xgboost-xgbDART) led to similar results, whereas in NN, the differences between the implementations used (nnet-nnet and nnet-avNNet) were more distinct. Finally, ML models tuned through a random search optimization method were compared with the same ML models with their default values. The results showed that the predictions were improved by the optimization process only where the ML algorithms demanded a large number of hyperparameters that needed tuning and there was a significant difference between the default values and the optimized ones, like in the case of GB and NN, but not in RF. In general, the current study concluded that although RF and GB presented approximately the same prediction accuracy, RF had more consistent results, regardless of different packages, different hyperparameter selection methods, or even the inclusion of OK in the ML models’ residuals.

  • PDF Download Icon
  • Research Article
  • Cite Count Icon 9
  • 10.3390/bioengineering10080967
Non-Contrasted CT Radiomics for SAH Prognosis Prediction
  • Aug 16, 2023
  • Bioengineering
  • Dezhi Shan + 4 more

Subarachnoid hemorrhage (SAH) denotes a serious type of hemorrhagic stroke that often leads to a poor prognosis and poses a significant socioeconomic burden. Timely assessment of the prognosis of SAH patients is of paramount clinical importance for medical decision making. Currently, clinical prognosis evaluation heavily relies on patients’ clinical information, which suffers from limited accuracy. Non-contrast computed tomography (NCCT) is the primary diagnostic tool for SAH. Radiomics, an emerging technology, involves extracting quantitative radiomics features from medical images to serve as diagnostic markers. However, there is a scarcity of studies exploring the prognostic prediction of SAH using NCCT radiomics features. The objective of this study is to utilize machine learning (ML) algorithms that leverage NCCT radiomics features for the prognostic prediction of SAH. Retrospectively, we collected NCCT and clinical data of SAH patients treated at Beijing Hospital between May 2012 and November 2022. The modified Rankin Scale (mRS) was utilized to assess the prognosis of patients with SAH at the 3-month mark after the SAH event. Based on follow-up data, patients were classified into two groups: good outcome (mRS ≤ 2) and poor outcome (mRS > 2) groups. The region of interest in NCCT images was delineated using 3D Slicer software, and radiomic features were extracted. The most stable and significant radiomic features were identified using the intraclass correlation coefficient, t-test, and least absolute shrinkage and selection operator (LASSO) regression. The data were randomly divided into training and testing cohorts in a 7:3 ratio. Various ML algorithms were utilized to construct predictive models, encompassing logistic regression (LR), support vector machine (SVM), random forest (RF), light gradient boosting machine (LGBM), adaptive boosting (AdaBoost), extreme gradient boosting (XGBoost), and multi-layer perceptron (MLP). Seven prediction models based on radiomic features related to the outcome of SAH patients were constructed using the training cohort. Internal validation was performed using five-fold cross-validation in the entire training cohort. The receiver operating characteristic curve, accuracy, precision, recall, and f-1 score evaluation metrics were employed to assess the performance of the classifier in the overall dataset. Furthermore, decision curve analysis was conducted to evaluate model effectiveness. The study included 105 SAH patients. A comprehensive set of 1316 radiomics characteristics were initially derived, from which 13 distinct features were chosen for the construction of the ML model. Significant differences in age were observed between patients with good and poor outcomes. Among the seven constructed models, model_SVM exhibited optimal outcomes during a five-fold cross-validation assessment, with an average area under the curve (AUC) of 0.98 (standard deviation: 0.01) and 0.88 (standard deviation: 0.08) on the training and testing cohorts, respectively. In the overall dataset, model_SVM achieved an accuracy, precision, recall, f-1 score, and AUC of 0.88, 0.84, 0.87, 0.84, and 0.82, respectively, in the testing cohort. Radiomics features associated with the outcome of SAH patients were successfully obtained, and seven ML models were constructed. Model_SVM exhibited the best predictive performance. The radiomics model has the potential to provide guidance for SAH prognosis prediction and treatment guidance.

  • Research Article
  • Cite Count Icon 41
  • 10.1186/s12877-022-03502-9
Application of machine learning model to predict osteoporosis based on abdominal computed tomography images of the psoas muscle: a retrospective study
  • Oct 13, 2022
  • BMC Geriatrics
  • Cheng-Bin Huang + 5 more

BackgroundWith rapid economic development, the world's average life expectancy is increasing, leading to the increasing prevalence of osteoporosis worldwide. However, due to the complexity and high cost of dual-energy x-ray absorptiometry (DXA) examination, DXA has not been widely used to diagnose osteoporosis. In addition, studies have shown that the psoas index measured at the third lumbar spine (L3) level is closely related to bone mineral density (BMD) and has an excellent predictive effect on osteoporosis. Therefore, this study developed a variety of machine learning (ML) models based on psoas muscle tissue at the L3 level of unenhanced abdominal computed tomography (CT) to predict osteoporosis.MethodsMedical professionals collected the CT images and the clinical characteristics data of patients over 40 years old who underwent DXA and abdominal CT examination in the Second Affiliated Hospital of Wenzhou Medical University database from January 2017 to January 2021. Using 3D Slicer software based on horizontal CT images of the L3, the specialist delineated three layers of the region of interest (ROI) along the bilateral psoas muscle edges. The PyRadiomics package in Python was used to extract the features of ROI. Then Mann–Whitney U test and the least absolute shrinkage and selection operator (LASSO) algorithm were used to reduce the dimension of the extracted features. Finally, six machine learning models, Gaussian naïve Bayes (GNB), random forest (RF), logistic regression (LR), support vector machines (SVM), Gradient boosting machine (GBM), and Extreme gradient boosting (XGBoost), were applied to train and validate these features to predict osteoporosis.ResultsA total of 172 participants met the inclusion and exclusion criteria for the study. 82 participants were enrolled in the osteoporosis group, and 90 were in the non-osteoporosis group. Moreover, the two groups had no significant differences in age, BMI, sex, smoking, drinking, hypertension, and diabetes. Besides, 826 radiomic features were obtained from unenhanced abdominal CT images of osteoporotic and non-osteoporotic patients. Five hundred fifty radiomic features were screened out of 826 by the Mann–Whitney U test. Finally, 16 significant radiomic features were obtained by the LASSO algorithm. These 16 radiomic features were incorporated into six traditional machine learning models (GBM, GNB, LR, RF, SVM, and XGB). All six machine learning models could predict osteoporosis well in the validation set, with the area under the receiver operating characteristic (AUROC) values greater than or equal to 0.8. GBM is more effective in predicting osteoporosis, whose AUROC was 0.86, sensitivity 0.70, specificity 0.92, and accuracy 0.81 in validation sets.ConclusionWe developed six machine learning models to predict osteoporosis based on psoas muscle images of abdominal CT, and the GBM model had the best predictive performance. GBM model can better help clinicians to diagnose osteoporosis and provide timely anti-osteoporosis treatment for patients. In the future, the research team will strive to include participants from multiple institutions to conduct external validation of the ML model of this study.

  • PDF Download Icon
  • Research Article
  • Cite Count Icon 20
  • 10.1155/2022/8089428
Predicting and Investigating the Permeability Coefficient of Soil with Aided Single Machine Learning Algorithm
  • Jan 1, 2022
  • Complexity
  • Van Quan Tran

The permeability coefficient of soils is an essential measure for designing geotechnical construction. The aim of this paper was to select a highest performance and reliable machine learning (ML) model to predict the permeability coefficient of soil and quantify the feature importance on the predicted value of the soil permeability coefficient with aided machine learning‐based SHapley Additive exPlanations (SHAP) and Partial Dependence Plot 1D (PDP 1D). To acquire this purpose, five single ML algorithms including K‐nearest neighbors (KNN), support vector machine (SVM), light gradient boosting machine (LightGBM), random forest (RF), and gradient boosting (GB) are used to build ML models for predicting the permeability coefficient of soils. Performance criteria for ML models include the coefficient of correlation R 2 , root mean square error (RMSE), mean absolute percentage error (MAPE), and mean absolute error (MAE). The best performance and reliable single ML model for predicting the permeability coefficient of soil for the testing dataset is the gradient boosting (GB) model, which has R 2 = 0.971, RMSE = 0.199 × 10 −11 m/s, MAE = 0.161 × 10 −11 m/s, and MAPE = 0.185%. To identify and quantify the feature importance on the permeability coefficient of soil, sensitivity studies using permutation importance, SHapley Additive exPlanations (SHAP), and Partial Dependence Plot 1D (PDP 1D) are performed with the aided best performance and reliable ML model GB. Plasticity index, density > water content, liquid limit, and plastic limit > clay content > void ratio are the order effects on the predicted value of the permeability coefficient. The plasticity index and density of soil are the first priority soil properties to measure when assessing the permeability coefficient of soil.

  • Research Article
  • Cite Count Icon 1
  • 10.3760/cma.j.cn112137-20221018-02174
The value of machine learning models based on biparametric MRI for diagnosis of prostate cancer and clinically significant prostate cancer
  • May 23, 2023
  • Zhonghua yi xue za zhi
  • X M Wang + 7 more

Objective: To evaluate the value of machine learning (ML) models based on biparametric magnetic resonance imaging (bpMRI) for diagnosis of prostate cancer (PCa) and clinically significant prostate cancer (csPCa). Methods: A total of 1 368 patients, aged from 30 to 92 (69.4±8.2) years, from 3 tertiary medical centers in Jiangsu Province were retrospectively collected from May 2015 to December 2020, including 412 cases of csPCa, 242 cases of clinically insignificant prostate cancer (ciPCa) and 714 cases of benign prostate lesions. The data of center 1 and center 2 were randomly divided into training cohort and internal testing cohort at a ratio of 7∶3 by random number sampling without replacement using Python Random package, and the data of center 3 were used as the independent external testing cohort. The training cohort includs 243 cases of csPCa, 135 cases of ciPCa and 384 cases of benign lesions, the internal testing cohort includs 104 cases of csPCa, 58 cases of ciPCa and 165 cases of benign lesions, and the external testing cohort includs 65 cases of csPCa, 49 cases of ciPCa and 165 cases of benign lesions. The radiomics features were extracted on T2-weighted imaging, diffusion-weighted imaging and apparent diffusion coefficient map, and optimal radiomics features were selected by using Pearson correlation coefficient method and analysis of variance. The ML models were built using two ML algorithms, including support vector machine and random forest (RF) and were further tested in the internal testing cohort and external testing cohort. Finally, the PI-RADS scores evaluated by the radiologists were adjusted by the ML models which had superior diagnostic performance, namely adjusted PI-RADS. The receiver operating characteristic (ROC) curves were used to evaluate the diagnostic performance of the ML models and PI-RADS. DeLong test was used to compare the areas under curve (AUC) of models with those of PI-RADS. Results: For PCa diagnosis, in internal testing cohort, the AUC of ML model using RF algorithm and PI-RADS were 0.869 (95%CI: 0.830-0.908) and 0.874 (95%CI: 0.836-0.913), respectively, and the difference between the model and PI-RADS did not reach to the statistical significance (P=0.793). In the external testing cohort, the AUC of model and PI-RADS were 0.845 (95%CI: 0.794-0.897) and 0.915 (95%CI: 0.880-0.951), respectively, and the difference was statistically significant (P=0.01). For csPCa diagnosis, the AUC of ML model using RF algorithm and PI-RADS were 0.874 (95%CI: 0.834-0.914) and 0.892 (95%CI: 0.857-0.927), respectively, in internal testing cohort, and the difference between the model and PI-RADS was not statistically significant (P=0.341). In the external testing cohort, the AUC of model and PI-RADS were 0.876 (95%CI: 0.831-0.920) and 0.884 (95%CI: 0.841-0.926), respectively, and the difference between the model and PI-RADS was not statistically significant (P=0.704). When PI-RADS assessment was adjusted with the assistance of ML models, the specificities increased from 63.0% to 80.0% in the internal testing cohort and from 92.7% to 93.3% in the external test group in diagnosing PCa. In diagnosing csPCa, the specificities increased from 52.5% to 72.6% in the internal testing cohort and from 75.2% to 79.9% in the external testing cohort. Conclusions: The ML models based on bpMRI showed comparable diagnostic performance to PI-RADS assessed by senior radiologists and achieved good generalization ability in both diagnosing PCa and csPCa. The specificities of the PI-RADS were improved by ML models.

  • Research Article
  • 10.21037/jtd-2026-1-0339
FDG-PET/CT data enhances machine learning prediction of occult lymph-node metastasis in lung adenocarcinoma
  • Apr 24, 2026
  • Journal of Thoracic Disease
  • Akihiro Sasaki + 11 more

BackgroundOccult lymph-node metastasis (ONM) occurs in 10–20% of primary lung cancers and can influence surgical and treatment strategies. Improving preoperative prediction of ONM may optimize lymph-node dissection and multidisciplinary planning. We aimed to develop machine learning (ML) models to predict ONM using clinical data and to assess whether adding fluorodeoxyglucose positron emission tomography (FDG-PET)/computed tomography (CT) data, including individual lymph-node maximum standardized uptake value (SUVmax), could enhance predictive performance.MethodsWe retrospectively analyzed 77 patients who underwent curative lobectomy for primary lung adenocarcinoma at our institution between 2016 and 2018. Four ML models—Random Forest (RF), Support Vector Machine (SVM), Gradient Boosting Machine (GBM), and eXtreme Gradient Boosting (XGB)—were developed using clinical data, PET/CT data, and their combination. Feature selection was performed using the Minimum Redundancy Maximum Relevance (mRMR) algorithm to reduce feature redundancy. The dataset was divided into training and testing sets in an 8:2 ratio with no overlapping. Hyperparameters were tuned using five-fold cross-validation. Model performance was evaluated using the area under the receiver operating characteristic curve (AUC) and average precision (AP), with confidence intervals estimated via bootstrap resampling. Feature importance was assessed using SHapley Additive exPlanations (SHAP).ResultsUsing clinical data alone, AUCs were 0.88 (RF), 0.69 (SVM), 0.73 (GBM), and 0.82 (XGB). Incorporation of PET/CT features improved predictive performance, yielding AUCs of 0.91 (SVM), 0.87 (GBM), and 0.91 (XGB). SHAP analysis demonstrated that while primary tumor PET features were important, the SUVmax of individual lymph nodes emerged as additional key predictors, highlighting the critical contribution of nodal PET data to ONM prediction.ConclusionsIncorporating PET/CT data, including individual lymph node SUVmax, improved ONM prediction in ML models. These findings highlight the critical contribution of nodal PET data and support the integration of imaging data into ML tools to guide personalized surgical and treatment decisions in lung adenocarcinoma.

  • Research Article
  • Cite Count Icon 3
  • 10.1007/s11604-023-01475-2
Utility of machine learning for identifying stapes fixation on ultra-high-resolution CT.
  • Aug 10, 2023
  • Japanese journal of radiology
  • Ruowei Tang + 8 more

Imaging diagnosis of stapes fixation (SF) is challenging owing to a lack of definite evidence. We developed a comprehensive machine learning (ML) model to identify SF on ultra-high-resolution CT. We retrospectively enrolled 109 participants (143 ears) and divided them into the training set (115 ears) and test set (28 ears). Stapes mobility (SF or non-SF) was determined by surgical inspection. In the ML analysis, rectangular regions of interest were placed on consecutive axial slices in the training set. Radiomic features were extracted and fed into the training session. The test set was analyzed using 7 ML models (support vector machine, k nearest neighbor, decision tree, random forest, extra trees, eXtreme Gradient Boosting, and Light Gradient Boosting Machine) and by 2 dedicated neuroradiologists. Diagnostic performance (sensitivity, specificity and accuracy, with surgical findings as the reference) was compared between the radiologists and the optimal ML model by using the McNemar test. The mean age of the participants was 42.3 ± 17.5years. The Light Gradient Boosting Machine (LightGBM) model showed the highest sensitivity (0.83), specificity (0.81), accuracy (0.82) and area under the curve (0.88) for detecting SF among the 7 ML models. The neuroradiologists achieved good sensitivities (0.75 and 0.67), moderate-to-good specificities (0.63 and 0.56) and good accuracies (0.68 and 0.61). This model showed no statistical differences with the neuroradiologists (P values 0.289-1.000). Compared to the neuroradiologists, the LightGBM model achieved competitive diagnostic performance in identifying SF, and has the potential to be a supportive tool in clinical practice.

  • Preprint Article
  • 10.5194/ems2025-562
Hydrological modelling using machine and deep learning models across multiple case studies
  • Jul 16, 2025
  • Majid Niazkar + 3 more

Machine learning (ML) and deep learning (DL) models can play an important role when it comes to modelling complicated processes. Such capability is necessary for hydrological and climate-related applications. Generally, ML models utilize precipitation and temperature time series of a basin as input to develop a lumped rainfall-runoff model to simulate streamflow at the basin outlet. However, when it is divided into several sub-basins, Graph Neural Networks (GNN) can consider each sub-basin as a node and link them together using a connectivity matrix to account for spatial variations of hydroclimatic variables. In this study, GNN and various ML models with different types of architecture, ranging from neural networks, tree-based structure, and gradient boosting, were exploited for daily streamflow simulation over different case studies. For each case study, the basin was divided into a few sub-basins for which daily precipitation and temperature data were aggregated and used as input. For training GNN, the connection matrix of sub-basins was also used as input. Basically, 75% of historical records were utilized to train GNN and different ML models, e.g., artificial neural networks, support vector machine, decision tree, random forest, eXtreme Gradient Boosting (XGBoost), Light Gradient-Boosting Machine (LightGBM), and Category Boosting (CatBoost), while the rest was used for testing. Streamflow simulation was conducted with/without considering seasonality impact and lag times. The obtained results clearly demonstrate that considering seasonality and time lags can enhance accuracy of streamflow predictions based on Kling–Gupta efficiency (KGE). Furthermore, GNN with seasonality impact and time lags achieved promising results across different case studies with KGE>0.85 for training and KGE>0.59 for testing data, respectively. Among ML models, boosting models, e.g., LightGBM and XGBoost, performed slightly better than other ML models. for Finally, this comparative analysis provides valuable insights for ML/DL applications in climate change impact assessments.Acknowledgements: This research work was carried out as part of the TRANSCEND project with funding received from the European Union Horizon Europe Research and Innovation Programme under Grant Agreement No. 10108411.

  • PDF Download Icon
  • Research Article
  • 10.3389/fonc.2026.1727595
Machine learning model for predicting malnutrition risk in lung cancer patients after thoracoscopic resection: a multi-center study.
  • Feb 9, 2026
  • Frontiers in oncology
  • Tianfeng Chen + 6 more

Early detection of malnutrition is critical for timely intervention in lung cancer patients undergoing thoracoscopic resection. Existing black-box prediction models lack clinical interpretability, limiting trust and application. The present study was conducted to predict malnutrition risk by establishing an explainable machine learning (ML) model and evaluate the model performance across several sites, so as to develop a web-based application to aid clinical decision-making. A retrospective analysis was conducted on 1, 134 lung cancer patients who underwent thoracoscopic resection at Dongguan People's Hospital between October 2021 and October 2024, consisting of a training set (n = 795) and a testing set (n = 339). Meanwhile, an external validation cohort (n=273) was prospectively enrolled at the Affiliated Hospital of Guangdong Medical University from March to June of 2025. Furthermore, univariate and multivariate analyses were employed to determine the individual risk variables for post-operative malnutrition. This study constructed eight ML models using Gradient Boosting Machine (GBM), Neural Network, Logistic Regression, Extreme Gradient Boosting (XGBoost), Random Forest, K-Nearest Neighbors (KNN), Adaptive Boosting (AdaBoost), and Support Vector Machine (SVM). The performance of the established models was assessed by decision curve analysis (DCA) and receiver operating characteristic (ROC) curves. Meanwhile, feature contributions and visualize model outputs were quantified using the SHapley Additive exPlanations (SHAP) method to enhance clinical interpretability. Consequently, a web-based risk calculator was created to assist in personalized forecasting. Among 1, 407 total patients, post-operative malnutrition incidence was 11.3% (159/1, 407). Multivariate analysis identified seven independent risk factors: albumin (ALB), Nutritional Risk Screening 2002 score, age, intraoperative blood loss, total drainage volume, Basic Activities of Daily Living (BADL) score, and serum potassium (K). The XGBoost model outperformed others, with AUC 0.845 (95% CI: 0.771-0.919) in the testing set and 0.886 (95% CI: 0.841-0.932) in external validation. SHAP analysis clarified the relative importance of risk factors, improving interpretability. The XGBoost-based explainable ML model effectively predicts malnutrition risk in lung cancer patients after thoracoscopic resection. Integrating high predictive performance with interpretability, it supports clinical risk stratification and personalized nutritional interventions to improve post-operative outcomes. A publicly available web-based calculator facilitates easy clinical application.

  • Research Article
  • 10.21037/tcr-2025-1-2845
Magnetic resonance imaging-based radiomics analysis: predicting vascular invasion in breast invasive ductal carcinoma using different machine learning models
  • Mar 26, 2026
  • Translational Cancer Research
  • Hongen Li + 4 more

BackgroundLymphovascular invasion (LVI) is an adverse prognostic factor; preoperative prediction by imaging is difficult, but radiomics extracts quantitative tumor biology features. Investigating the value of combining magnetic resonance imaging (MRI)-based radiomics features with multiple machine learning (ML) models in predicting LVI status in invasive ductal carcinoma (IDC) of the breast.MethodsA retrospective cohort of 678 female patients with pathologically confirmed IDC of the breast was collected from June 2021 to June 2025. All patients underwent preoperative MRI. Based on postoperative pathology, patients were categorized into LVI-positive (n=258) and LVI-negative (n=420) groups. Using ITK-SNAP software, regions of interest (ROIs) were delineated in phase-3 dynamic contrast-enhanced MRI images to extract radiomics features. Feature selection and dimensionality reduction were performed using redundancy analysis and the least absolute shrinkage and selection operator (LASSO) regression. Data were randomly split into an 8:2 ratio for training (n=542) and testing (n=136) sets. Eight ML models were then constructed: logistic regression (LR), support vector machine (SVM), K-nearest neighbors (KNN), random forest, extreme random trees (ExtraTrees), extreme gradient boosting (XGBoost), light gradient boosting machine (LightGBM), and multi-layer perceptron (MLP). Univariate and multivariate LR analyses were performed to screen clinical and radiological features for establishing clinical models. Concurrently, a combined model integrating radiomics features with clinical characteristics was developed. The discriminatory power of each model was evaluated using the area under the curve (AUC). AUC values for the radiological model, clinical model, and combined model underwent statistical comparison via Delong’s test. Decision curve analysis (DCA) was employed to assess their clinical utility.ResultsA total of 1,197 radiomics features were extracted, and after dimensionality reduction, 23 features with the highest predictive value were selected. The clinical prediction model constructed based on multifactorial analysis results indicated that LVI positivity was more likely to occur in postmenopausal patients [odds ratio (OR) =1.690; 95% confidence interval (CI): 1.174–2.433], those with higher histological grade (OR =1.527; 95% CI: 1.107–2.107), sentinel lymph node metastasis (OR =0.198; 95% CI: 0.137–0.285), distinct molecular subtypes (OR =0.740; 95% CI: 0.567–0.965), and MRI maximum diameter ≥2 cm (OR =2.059; 95% CI: 1.362–3.113). Among radiomics models, the XGBoost model demonstrated optimal performance with a training set AUC of 0.912 and a validation set AUC of 0.706. The combined model exhibited the highest discriminatory ability in the training set (AUC =0.956) and a validation set AUC of 0.778. DCA indicated the combined model provided higher clinical net benefit.ConclusionsA combined model incorporating MRI radiomics features and clinical factors demonstrates predictive value for the presence or absence of LVI in IDC of the breast, serving as a reference for individualized treatment decisions.

  • Research Article
  • Cite Count Icon 31
  • 10.1016/j.eswa.2023.119768
Ensemble machine learning-based models for estimating the transfer length of strands in PSC beams
  • Mar 1, 2023
  • Expert Systems with Applications
  • Viet-Linh Tran + 1 more

Ensemble machine learning-based models for estimating the transfer length of strands in PSC beams

  • Research Article
  • 10.55041/ijsrem32427
Performance Comparison of Various Machine Learning Models for Rainfall Prediction Using Wind Information
  • May 4, 2024
  • INTERANTIONAL JOURNAL OF SCIENTIFIC RESEARCH IN ENGINEERING AND MANAGEMENT
  • Vishal Jain

Forecasting rainfall, crucial for agriculture, water management, and disaster preparedness, presents significant challenges due to intricate relationships often missed by conventional statistical methods. Hence, machine learning (ML) models offer promising alternatives, enhancing the precision and dependability of rainfall predictions. This paper presents a comprehensive comparison of various ML models with diverse model structures and regularization strategies for rainfall prediction in urban metropolitan cities. The results show that the random forest model, and gradient boosting model outperform the other models such as logistic regression, support vector machine (SVM), decision tree, K-nearest neighbor (KNN), Naive Bayes, linear SVM, and neural network in terms of accuracy. Validation accuracies of 75%, 77%, 68%, 78%, 76%, 78%, 74%, 75% and 76% were achieved for logistic regression, SVM, decision tree, random forest model, KNN, gradient boosting model, Naive Bayes, linear SVM, and neural network, respectively. The choice of ML models for rainfall prediction should consider the characteristics of the data, e.g. a lag feature for 20 days was employed that uses previous time steps to predict the next time step. The paper concludes that ML models, especially the random forest and gradient boosting models are powerful and robust tools for rainfall prediction in urban metropolitan cities. Key Words: rainfall prediction, wind information, machine learning, gradient boosting, random forest

  • Research Article
  • Cite Count Icon 27
  • 10.1016/j.geoen.2023.212086
Machine learning approaches for formation matrix volume prediction from well logs: Insights and lessons learned
  • Jul 8, 2023
  • Geoenergy Science and Engineering
  • Pamidi Venkata Durga Kannaiah + 1 more

Machine learning approaches for formation matrix volume prediction from well logs: Insights and lessons learned

  • Research Article
  • Cite Count Icon 8
  • 10.1136/postgradmedj-2021-141329
Machine learning algorithm can provide assistance for the diagnosis of non-ST-segment elevation myocardial infarction
  • Feb 16, 2022
  • Postgraduate Medical Journal
  • Lian Qin + 5 more

IntroductionOur aim was to use the constructed machine learning (ML) models as auxiliary diagnostic tools to improve the diagnostic accuracy of non-ST-elevation myocardial infarction (NSTEMI).Materials and methodsA total of 2878...

Save Icon
Up Arrow
Open/Close
Notes

Save Important notes in documents

Highlight text to save as a note, or write notes directly

You can also access these Documents in Paperpal, our AI writing tool

Powered by our AI Writing Assistant