Accelerate Literature Icon
Want to do a literature review? Try our new Literature Review workflow

Machine Learning Algorithms to Detect Sex in Myocardial Perfusion Imaging.

  • Abstract
  • Highlights & Summary
  • Literature Map
  • Similar Papers
Abstract
Translate article icon Translate Article Star icon

Myocardial perfusion imaging (MPI) is an essential tool used to diagnose and manage patients with suspected or known coronary artery disease. Additionally, the General Data Protection Regulation (GDPR) represents a milestone about individuals' data security concerns. On the other hand, Machine Learning (ML) has had several applications in the most diverse knowledge areas. It is conceived as a technology with huge potential to revolutionize health care. In this context, we developed ML models to evaluate their ability to distinguish an individual's sex from MPI assessment. We used 260 polar maps (140 men/120 women) to train ML algorithms from a database of patients referred to a university hospital for clinically indicated MPI from January 2016 to December 2018. We tested 07 different ML models, namely, Classification and Regression Tree (CART), Naive Bayes (NB), K-Nearest Neighbors (KNN), Support Vector Machine (SVM), Adaptive Boosting (AB), Random Forests (RF) and, Gradient Boosting (GB). We used a cross-validation strategy. Our work demonstrated that ML algorithms could perform well in assessing the sex of patients undergoing myocardial scintigraphy exams. All the models had accuracy greater than 82%. However, only SVM achieved 90%. KNN, RF, AB, GB had, respectively, 88, 86, 85, 83%. Accuracy standard deviation was lower in KNN, AB, and RF (0.06). SVM and RF had had the best area under the receiver operating characteristic curve (0.93), followed by GB (0.92), KNN (0.91), AB, and NB (0.9). SVM and AB achieved the best precision. Our results bring some challenges regarding the autonomy of patients who wish to keep sex information confidential and certainly add greater complexity to the debate about what data should be considered sensitive to the light of the GDPR.

Similar Papers
  • Research Article
  • 10.1002/cpe.70325
Soil Nutrient Analysis and Yield Prediction With Neuro‐ ML Ensemble Model Using IoT ‐ WSN Approach: In Context to India's Agricultural Sector
  • Oct 21, 2025
  • Concurrency and Computation: Practice and Experience
  • Sandeep Bhatia + 2 more

Agriculture is a backbone of the Indian economy and people's lives. In agriculture land, soil is the most important element on which the quality of production and efficiency depends to the maximum extent. Phosphorus (P), Nitrogen (N), Potassium (K), and the potential of hydrogen (pH) are the key nutrients in soil. An efficient crop recommender and prediction system is needed to optimize agriculture practices considering the escalating demand for more food. Traditional time‐consuming and manual farming should be replaced with a smart agriculture framework using the integration of technologies like the Internet of Things (IoT), Wireless Sensor Network (WSN), and Machine Learning (ML). This paper proposed an IoT‐WSN driven crop management system with Neuro‐ML Ensemble Model, utilizing LoRaWAN Gateway, that can be deployed in the agriculture field to collect real‐time soil parameters. In this paper for soil nutrient analysis, the author used various ML algorithms such as Naive Bayes (NB), Logistic Regression (LR), K‐Nearest Neighbor (KNN), Decision Tree (DT), Random Forest (RF), Ada Boost (AB), Gradient Boosting (GB), and Support Vector Machine (SVM) and recommending a suitable ML algorithm for the crop recommender system. For crop yield prediction, the author has developed and recommended a customized GB Algorithm with an accuracy of 98.80%, and for the fertilizer recommendation system, the author has suggested CNN‐BiGRU which outperforms other approaches like BiGRU and CNN with an average accuracy rate of 92.48%. The author presented work with respect to the Indian agriculture sector and compared ML algorithms with state‐of‐the‐art datasets available on some government websites of India, and used by other authors, with a dataset collected by the author from hardware using Raspberry Pi. For crop recommendation and forecasting, the Neuro‐ML Ensemble model employs the Neuro‐ML, which combines neural networks (NN) with the ML models. This research aspires to assist farmers in opting for suitable crops as per their environmental suitability and situation by analyzing and predicting which crops suit well to fit the parameters required to enhance crop growth like soil nutrients, soil moisture, soil pH, and rainfall, etc. The author obtained accuracy for various ML models used in the framework. For NB, LR, KNN, SVM, DT, and RF, the author obtained accuracies of 99.54%, 96.36%, 95.90%, 96.81%, 98.86%, and 99.31%, respectively, using the Kaggle dataset available as open access. Through a dataset collected by the authors, we obtained accuracies of 94.54%, 91.36%, 92.72%, 92.73%, 86.36%, and 94.54% for NB, LR, KNN, SVM, DT, and RF, respectively. The author found that Naive Bayes (NB) outperforms the other machine learning algorithms, such as KNN, SVM, LR, Decision Tree, RF, and AB, and is the best algorithm suited for crop yield.

  • Research Article
  • Cite Count Icon 13
  • 10.1016/j.tws.2024.112427
Regression-classification ensemble machine learning model for loading capacity and bucking mode prediction of cold-formed steel built-up I-section columns
  • Sep 7, 2024
  • Thin-Walled Structures
  • Yan Lu + 4 more

Regression-classification ensemble machine learning model for loading capacity and bucking mode prediction of cold-formed steel built-up I-section columns

  • Research Article
  • Cite Count Icon 15
  • 10.1007/s10753-023-01827-0
Performance of Machine Learning Algorithms for Predicting Disease Activity in Inflammatory Bowel Disease.
  • May 12, 2023
  • Inflammation
  • Weimin Cai + 5 more

This study aimed to explore the effectiveness of predicting disease activity in patients with inflammatory bowel disease (IBD), using machine learning (ML) models. A retrospective research was undertaken on IBD patients who were admitted intothe First Affiliated Hospital of Wenzhou Medical University between September 2011 and September 2019. At first, data were randomly split into a 3:1 ratio of training to test set. The least absolute shrinkage and selection operator (LASSO) algorithm was applied to reduce the dimension of variables. These variables were used to generate seven ML algorithms, namely random forests (RFs), adaptive boosting (AdaBoost), K-nearest neighbors (KNNs), support vector machines (SVMs), naïve Bayes (NB), ridge regression, and eXtreme gradient boosting (XGBoost) to train to predict disease activity in IBD patients. SHapley Additive exPlanation (SHAP) analysis was performed to rank variable importance. A total of 876 participants with IBD, consisting of 275 ulcerative colitis (UC) and 601 Crohn's disease (CD), were retrospectively enrolled in the study. Thirty-three variables were obtained from the clinical characteristics and laboratory tests of the participants. Finally, after LASSO analysis, 11 and 5 variables were screened out to construct ML models for CD and UC, respectively. All seven ML models performed well in predicting disease activity in the CD and UC test sets. Among these ML models, SVM was more effective in predicting disease activity in the CD group, whose AUC reached 0.975, sensitivity 0.947, specificity 0.920, and accuracy 0.933. AdaBoost performed best for the UC group, with an AUC of 0.911, sensitivity 0.844, specificity 0.875, and accuracy 0.855. ML algorithms were available and capable of predicting disease activity in IBD patients. Based on clinical and laboratory variables, ML algorithms demonstrate great promise in guiding physicians' decision-making.

  • Conference Article
  • Cite Count Icon 10
  • 10.2118/218838-ms
Machine Learning Models to Predict Total Skin Factor in Perforated Wells
  • Apr 9, 2024
  • SPE Western Regional Meeting
  • S Thabet + 6 more

An accurate total skin factor prediction for an oil well is critical for the evaluation of the inflow performance relationship, and the optimization of the appropriate stimulation treatment such as acidizing and hydraulic fracturing. Performing well testing regularly is not economically feasible, and the equations used for total skin damage may not be accurate. In this work, the goal is to build machine learning (ML) models that can predict the total skin factor in perforated wells using accessible field data. Nine distinct ML algorithms such as Gradient Boosting (GB), Adaptive Boosting (AdaBoost), Random Forest (RF), Support Vector Machines (SVMs), Decision Trees (DT), K-Nearest Neighbor (KNN), Linear Regression (LR), Stochastic Gradient Descent (SGD), and Artificial Neural Network (ANN) are meticulously developed and fine-tuned using a substantial dataset derived from 1,088 wells. The dataset encompasses 19,040 data points, thoughtfully split into two subsets: 70% (13,328 data points) for training the algorithms, and 30% (5,712 data points) for testing their predictions. The parameters used are mostly gathered during well completion and conventional well testing operations, including liquid flow rate, water cut, gas oil ratio, bottomhole flowing pressure, reservoir pressure, reservoir temperature, reservoir permeability, reservoir thickness, perforations diameter, perforations density, perforations penetration depth, well deviation, and penetrated portion of the net pay thickness. In this study, the total skin factor acquired from conventional well test analysis serves as the model's output. K-fold cross-validation and repeated random sampling validation techniques are used to assess the performance of the models against the total skin obtained from the conventional well test analysis. The K-fold cross-validation outcomes of the top-performing ML models, specifically GB, AdaBoost, RF, DT, and KNN, reveal remarkably low mean absolute percentage error values reported as 3.2%, 3.2%, 2.9%, 3.3%, and 3.8%, respectively. Additionally, the correlation coefficients (R2) for these models are notably high, with values of 0.972, 0.968, 0.975, 0.964, and 0.956, respectively. In conclusion, ML models demonstrated their ability to predict total skin factor for different reservoir fluid properties, well geometries, and completion configurations. ML models offer a more efficient, quick, and cost-effective alternative to the conventional well-testing analysis.

  • PDF Download Icon
  • Research Article
  • 10.1007/s44444-025-00045-3
Development of machine learning models for predicting the deposition of sulfide scales in oil production wells
  • Oct 28, 2025
  • Journal of King Saud University – Engineering Sciences
  • Mohamed Mostafa Askar + 3 more

Scale deposition in oil wells poses numerous operational challenges that can lead to blockage of the completion string; therefore, it is recommended to predict scale accumulation before its occurrence. The current study aims to develop Artificial Neural Network (ANN) models and other Ensemble Machine Learning (ML) models to detect the presence of sulfide scales in oil wells and estimate the percentage of the sulfide scale composition in the accumulated scale. The studied sulfide scales are zinc, lead, and iron sulfides. The investigated ML models are Artificial Neural Network, Decision Tree, Random Forest, Support Vector Machine, K-Nearest Neighbors, Gradient Boosting, Extreme Gradient Boosting, Adaptive Boosting, Light Gradient Boosting Machine, and Ridge Regression. The ML models were constructed using actual field data for lab-analyzed and physically collected scale samples. These samples were collected from 347 production wells in 17 fields over a 22-year period. The original database consists of 1486 scale deposition incidences, and it was randomly split into 80% for training and validation, and 20% for testing. The assigned input data for scale prediction involve the relevant surface and downhole data, including: water ion analysis, water pH, production rates of well fluids, and content of acid gases, downhole pressure, downhole temperature, and injection gas rate for gas-lifted wells. The ANN models were implemented in several field applications and yielded promising results compared to scale tendency calculations from commercial scale prediction software. Moreover, variant models of Gradient Boosting and KNN models showed the highest model accuracy in predicting sulfide scale percentage.

  • Research Article
  • 10.1161/circ.152.suppl_3.4361080
Abstract 4361080: Daily Dietary Nutrients Predict Progression of Cardiovascular-Kidney-Metabolic Syndrome: A Machine Learning and SHAP Interpretation Study
  • Nov 4, 2025
  • Circulation
  • Yiqun Miao + 2 more

Background: The Cardiovascular–Kidney–Metabolic (CKM) syndrome, recently proposed by the American Heart Association, underscores the interplay among metabolic, renal, and cardiovascular conditions. Early risk identification is essential for effective prevention. Although diet is central to metabolic health, the impact of daily nutrient intake on CKM progression remains unclear. Objective: The aim of this study was to develop and validate a machine learning (ML) model for predicting the progression of CKM syndrome from early (Stages 1-3) to advanced stage (Stage 4). Methods: The National Health and Nutrition Examination Survey 2005-2018 dataset was used for the analysis. Daily dietary nutrients were selected as primary features, with demographic and lifestyle factors incorporated to improve model performance. Feature preprocessing involved VIF-based removal of multicollinearity, class balancing via Synthetic Minority Over-sampling Technique (SMOTE), and predictor selection using the Boruta algorithm. Subsequently, six ML algorithms, namely Random Forest (RF), light gradient boosting machine (LightGBM), Naive Bayes (NB), Support Vector Machine (SVM), eXtreme Gradient Boost (XGBoost), and K-Nearest Neighbors (KNN) were employed to train ML models using 10-fold cross-validation. Results: A total of 12,376 participants were enrolled in the study. The negative Weighted Quantile Sum regression index demonstrated a statistically significant inverse association between the dietary nutrient mixture and CKM syndrome progression (OR = 0.68, 95% CI: 0.62-0.74, P < 0.01). After excluding multicollinear variables and selecting important predictors, the ML model retained 29 daily dietary nutrients features and 5 baseline characteristics. The RF model demonstrated superior performance compared to alternative ML algorithms, achieving an accuracy of 91.1%, a sensitivity of 93.7%, a specificity of 87.4%, an F1 score of 92.6%, and an AUC of 0.971[95%CI(0.969-0.974)]. SHAP analysis indicated that among the dietary variables, niacin, copper, and vitamin E were identified as the most important nutritional predictors. Among demographic features, age and sex were the most influential factors. Conclusions: RF exhibited the best performance for predicting the progression of CKM syndrome from early to advanced stages. SHAP value interpretation revealed that niacin played a dominant role in prediction, with copper, vitamin E, age and sex also emerging as key contributing factors.

  • Research Article
  • Cite Count Icon 9
  • 10.2147/rmhp.s346856
Comparison Between Statistical Model and Machine Learning Methods for Predicting the Risk of Renal Function Decline Using Routine Clinical Data in Health Screening
  • Apr 26, 2022
  • Risk Management and Healthcare Policy
  • Xia Cao + 4 more

PurposeUsing machine learning method to predict and judge unknown data offers opportunity to improve accuracy by exploring complex interactions between risk factors. Therefore, we evaluate the performance of machine learning (ML) algorithms and to compare them with logistic regression for predicting the risk of renal function decline (RFD) using routine clinical data.Patients and MethodsThis retrospective cohort study includes datasets from 2166 subjects, aged 35–74 years old, provided by an adult health screening follow-up program between 2010 and 2020. Seven different ML models were considered – random forest, gradient boosting, multilayer perceptron, support vector machine, K-nearest neighbors, adaptive boosting, and decision tree - and were compared with standard logistic regression. There were 24 independent variables, and the baseline estimate glomerular filtration rate (eGFR) was used as the predictive variable.ResultsA total of 2166 participants (mean age 49.2±11.2 years old, 63.3% males) were enrolled and randomly divided into a training set (n=1732) and a test set (n=434). The area under receiver operating characteristic curve (AUROC) for detecting RFD corresponding to the different models were above 0.85 during the training phase. The gradient boosting algorithms exhibited the best average prediction accuracy (AUROC: 0.914) among all algorithms validated in this study. Based on AUROC, the ML algorithms improved the RFD prediction performance, compared to logistic regression model (AUROC:0.882), except the K-nearest neighbors and decision tree algorithms (AUROC:0.854 and 0.824, respectively). However, the improvement differences with logistic regression were small (less than 4%) and nonsignificant.ConclusionOur results indicate that the proposed health screening dataset-based RFD prediction model using ML algorithms is readily applicable, produces validated results. But logistic regression yields as good performance as ML models to predict the risk of RFD with simple clinical predictors.

  • Research Article
  • Cite Count Icon 4
  • 10.3389/fcvm.2024.1504957
Development and validation of a prediction model for coronary heart disease risk in depressed patients aged 20 years and older using machine learning algorithms
  • Jan 9, 2025
  • Frontiers in Cardiovascular Medicine
  • Yicheng Wang + 3 more

BackgroundDepression is being increasingly acknowledged as an important risk factor contributing to coronary heart disease (CHD). Currently, there is no predictive model specifically designed to evaluate the risk of coronary heart disease among individuals with depression. We aim to develop a machine learning (ML) model that will analyze risk factors and forecast the probability of coronary heart disease in individuals suffering from depression.MethodsThis research employed data from the National Health and Nutrition Examination Survey (NHANES) from 2007–2018, which included 2,085 individuals who had previously been diagnosed with depression. The population was randomly divided into a training set and a validation set, with an 8:2 ratio. Univariate and multivariate logistic regression analyses were employed to identify independent risk factors for coronary heart disease in individuals with depression. Eight machine learning algorithms were applied to the training set to construct the model, including logistic regression (LR), random forest (RF), gradient boosting machine (GBM), support vector machine (SVM), extreme gradient boosting (XGBoost), classification and regression tree (CART), k-nearest neighbors (KNN), and neural network (NNET). The validation set are used to evaluate the various performances of eight machine learning models. Several evaluation metrics were employed to assess and compare the performance of eight different machine learning models, aiming to identify the most effective algorithm for predicting coronary heart disease risk in individuals with depression. The evaluation metrics applied in this study included the area under the receiver operating characteristic (ROC) curve, calibration curve, Brier scores, decision curve analysis (DCA), and the precision-recall (PR) curve. And internally validated by the bootstrap method.ResultsUnivariate and multivariate logistic regression analyses identified age, chest pain status, history of myocardial infarction, serum triglyceride levels, and education level as independent predictors of coronary heart disease risk. Eight machine learning algorithms are applied to construct the models, among which the Random Forest model has the best performance, with an (Area Under Curve) AUC of 0.987 for the random forest model in the training set, and an AUC of 0.848 for the PR curve. In the validation set, the random forest model achieves an AUC of 0.996, and an AUC of 0.960 for the PR curve, which demonstrates an excellent discriminative ability. Calibration curves indicated high congruence between observed and predicted odds, with minimal Brier scores of 0.026 and 0.021 for the training, respectively, reinforcing the model's ability to discriminate. Set and validation set, respectively, reinforcing the model's predictive accuracy. DCA curves confirmed net benefits of the random forest model across. Furthermore, the AUC of the random forest model was 0.928 after internal validation by bootstrap method, indicating that its discriminative ability is good, and the model is useful for clinical assessment of the risk of coronary heart disease in depressed people.ConclusionThe random forest algorithm exhibited the best predictive performance, potentially aiding clinicians in assessing the risk probabilities of coronary heart disease within this population.

  • Research Article
  • Cite Count Icon 7
  • 10.1186/s13048-025-01654-x
Construction and evaluation of machine learning-based prediction model for live birth following fresh embryo transfer in IVF/ICSI patients with polycystic ovary syndrome
  • Apr 4, 2025
  • Journal of Ovarian Research
  • Suqin Zhu + 6 more

ObjectiveTo investigate the determinants affecting live birth outcomes in fresh embryo transfer among polycystic ovary syndrome (PCOS) patients using various machine learning (ML) algorithms and to construct predictive models, offering novel insights for enhancing live birth rates in this specific group.MethodsA sum of 1,062 fresh embryo transfer cycles involving PCOS patients were analyzed, with 466 resulting in live births. The dataset was split randomly into training and testing subsets at a 7:3 ratio. Least absolute shrinkage and selection operator and recursive feature elimination methods were utilized for feature selection within the training data. A grid search strategy identified the optimal parameters for seven ML models: decision tree (DT), K-nearest neighbors (KNN), light gradient boosting machine (LightGBM), naive Bayes model(NBM), random forest (RF), support vector machine (SVM) and extreme gradient boosting (XGBoost). The evaluation of model effectiveness incorporated diverse metrics, encompassing area under the curve (AUC), accuracy, positive predictive value, negative predictive value, F1 score, and Brier score. Calibration curves and decision curve analysis were employed to ascertain the optimal model. Furthermore, Shapley additive explanations were applied to elucidate the importance of predictor variables in the top-performing model.ResultsThe AUC values of DT, KNN, LightGBM, NBM, RF, SVM and XGBoost models in the training set were 0.813, 1.000, 0.724, 0.791, 1.000, 0.819 and 0.853, respectively. Corresponding values in the testing set were 0.773, 0.719, 0.705, 0.764, 0.794, 0.806 and 0.822. XGBoost emerged as the most effective ML model. SHAP analysis revealed that variables encompassing embryo transfer count, embryo type, maternal age, infertility duration, body mass index, serum testosterone (T) levels, and progesterone (P) levels on the day of human chorionic gonadotropin administration were pivotal predictors of live birth outcomes in individuals with PCOS receiving fresh embryo transfer.ConclusionThis study developed a live birth prediction model tailored for PCOS fresh embryo transfer cycles, leveraging ML algorithms to compare the efficacy of multiple models. The XGBoost model demonstrated superior predictive capacity, enabling prompt and precise identification of critical risk factors influencing live birth outcomes in PCOS patients. These findings offer actionable insights for clinical intervention, guiding strategies to improve pregnancy outcomes in this population.Clinical trial numberNot applicable.

  • PDF Download Icon
  • Research Article
  • Cite Count Icon 92
  • 10.1109/access.2019.2933670
Ensemble Learners of Multiple Deep CNNs for Pulmonary Nodules Classification Using CT Images
  • Jan 1, 2019
  • IEEE Access
  • Baihua Zhang + 6 more

Various deep convolutional neural networks (CNNs) have been used to distinguish between benign and malignant pulmonary nodules using CT images. However, single learner usually presents unsatisfied performance due to limited hypothesis space, or falling into local minima, or wrong selection of hypothesis space. To tackle these issues, we propose to build ensemble learners through fusing multiple deep CNN learners for pulmonary nodules classification. CT image patches of 743 nodules are extracted from LIDC-IDRI database and utilized. First, eight deep CNN learners with different architectures are trained and evaluated by 10-fold cross-validation. Each nodule has eight predictions from the eight primary learners. Second, we fuse these eight predictions by the strategies of majority voting (VOT), averaging (AVE), or machine learning. Specifically, different machine learning algorithms including K-Nearest-Neighbor (KNN), Support Vector Machines (SVM), Naive Bayes (NB), Decision Trees (DT), Multi-layer Perceptron (MLP), Random Forests (RF), Gradient Boosting Regression Trees (GBRT) and Adaptive Boosting (AdaBoost) are implemented. Moreover, the correlation coefficients between the predictions of 10 ensemble learners are calculated, and the hierarchical clustering dendrogram is drawn. It is found that the ensemble learners achieve higher prediction accuracy (84.0% vs 81.7%) than single CNN learner. The overlap ratio among the 10 ensemble learners is much higher than that of the 8 primary learners (62.9% vs 33.2%). In addition, it is shown that ensemble learners are roughly divided into three categories: the first (SVM, MLP, GBRT and RF) achieves the best performance; the second (VOT and AVE) is better than the third (AdaBoost, DT, NB and KNN). VOT and AVE yield higher recall than the machine learning algorithms. These results indicate that ensemble learners based on multiple CNN learners can achieve better performances for pulmonary nodules classification using CT images and that preferred fusion strategies include SVM, MLP, GBRT and RF.

  • PDF Download Icon
  • Research Article
  • Cite Count Icon 23
  • 10.3390/ijerph20010107
Development of Machine Learning Model for Prediction of Demolition Waste Generation Rate of Buildings in Redevelopment Areas.
  • Dec 21, 2022
  • International Journal of Environmental Research and Public Health
  • Gi-Wook Cha + 3 more

Owing to a rapid increase in waste, waste management has become essential, for which waste generation (WG) information has been effectively utilized. Various studies have recently focused on the development of reliable predictive models by applying artificial intelligence to the construction and prediction of WG information. In this study, research was conducted on the development of machine learning (ML) models for predicting the demolition waste generation rate (DWGR) of buildings in redevelopment areas in South Korea. Various ML algorithms (i.e., artificial neural network (ANN), K-nearest neighbors (KNN), linear regression (LR), random forest (RF), and support vector machine (SVM)) were applied to the development of an optimal predictive model, and the main hyper parameters (HPs) for each algorithm were optimized. The results suggest that ANN-ReLu (coefficient of determination (R2) 0.900, the ratio of percent deviation (RPD) 3.16), SVM-polynomial (R2 0.889, RPD 3.00), and ANN-logistic (R2 0.883, RPD 2.92) are the best ML models for predicting the DWGR. They showed average errors of 7.3%, 7.4%, and 7.5%, respectively, compared to the average observed values, confirming the accurate predictive performance, and in the uncertainty analysis, the d-factor of the models appeared less than 1, showing that the presented models are reliable. Through a comparison with ML algorithms and HPs applied in previous related studies, the results herein also showed that the selection of various ML algorithms and HPs is important in developing optimal ML models for WG management.

  • Research Article
  • Cite Count Icon 1
  • 10.1186/s12911-025-02987-9
Development of machine learning models to predict the risk of fungal infection following flexible ureteroscopy lithotripsy
  • Apr 10, 2025
  • BMC Medical Informatics and Decision Making
  • Haofang Zhang + 8 more

BackgroundThe flexible ureteroscopy lithotripsy (F-URL) is an important treatment for upper urinary tract stones. However, urolithiasis, surgical procedures, and catheter placement are risk factors for fungal infections. Our study aimed to construct a machine learning algorithm predictive model to predict the risk of fungal infection following F-URL.MethodsThis study retrospectively collected the clinical data of patients who underwent F-URL at the Second Affiliated Hospital of Zhengzhou University from January 2016 to March 2024. The patients were divided into a non-fungal infection group and a fungal infection group based on whether a fungal infection occurred within three months post-surgery. The patient data from January 2016 to December 2023 were used as training data, and the patient data from January 2024 to March 2024 were used as testing set. The training data was randomly divided into a training set and validation set at a ratio of 90:10. Use LASSO regression to screen clinical features based on the training set. Nine machine learning algorithms, Logistic Regression (LR), k-Nearest Neighbours (KNN), Support Vector Machines (SVM), Random Forest (RF), Categorical Boosting (CatBoost), eXtreme Gradient Boosting (XGBoost), Adaptive Boosting (AdaBoost), Gradient Boosting Machines (GBM), and Neural Network (NNet), were used to construct models. The performance of these nine models was evaluated and the best predictive model was selected based on the validation set, and evaluate the best predictive model’s generalization ability using the testing set. Visualize the constructed optimal machine learning model using the SHapley additive interpretation (SHAP) value method. SHAP force plots were established to show the application of the prediction model at the individual level.ResultsA total of 13 clinical features were used to construct predictive models: age, diabetes mellitus (DM), history of malignancy, being bedridden, admission white blood cells (WBC), preoperative ureteral stenting, operation time, postoperative fever, postoperative Neu, carbapenem antibiotics use, duration of antibiotic therapy, length of hospital stay (LOS), and postoperative stent duration. Comparing the performance of 9 prediction models, we found that the model constructed using XGBoost algorithm had the best performance. The model constructed using XGBoost algorithm shows good discrimination, generalization and clinical applicability in the testing set.ConclusionsThe XGBoost model developed in this study has good predictive ability and clinical applicability for evaluating the risk of fungal infection following F-URL.

  • Research Article
  • Cite Count Icon 1
  • 10.3389/fendo.2025.1514397
Predicting isolated impaired glucose tolerance without oral glucose tolerance test using machine learning in Chinese Han men.
  • Apr 24, 2025
  • Frontiers in endocrinology
  • Lin Wang + 9 more

Isolated Impaired Glucose Tolerance (I-IGT) represents a specific prediabetic state that typically requires a standardized oral glucose tolerance test (OGTT) for diagnosis. This study aims to predict glucose tolerance status in Chinese Han men at fasting state using machine learning (ML) models with demographic, anthropometric, and laboratory data. The study population consisted of 1,117 Chinese Han men aged 50-87 years. Baseline variables including age, fasting plasma glucose (FPG), high blood pressure (HBP), body mass index (BMI), waist to hip ratio (WHR), total cholesterol (TC), triglyceride (TG), high-density lipoprotein cholesterol (HDL-C), and low-density lipoprotein cholesterol (LDL-C) were collected from electronic medical records (EMRs) for machine learning model training and validation. Support Vector Machine (SVM), Decision Tree (DT), Random Forest (RF), Logistic Regression (LR), K-Nearest Neighbors (KNN), Naive Bayes (NB), Adaptive Boosting (AdaBoost) and Gradient Boosting Machines (GBM) were tested for machine learning model performance comparison. Model performance was evaluated using metrics including accuracy, recall, F1 score, positive predictive value (PPV), negative predictive value (NPV), and the area under the receiver operating characteristic curve (AUC). Shapley Additive Explanations (SHAP) and confusion matrix plots were used for model interpretation. The RF model demonstrated the best overall performance with a 96.7% accuracy, recall of 91.4%, F1 score of 95.7%, PPV of 99.1%, and NPV of 95.6%. The AUC values for the SVM, DT, RF, LR, KNN, NB, AdaBoost, and GBM models were 0.97, 0.92, 0.96, 0.97, 0.88, 0.88, 0.97, and 0.97, respectively. While the RF model showed strong overall performance, the LR model had the highest AUC, indicating superior discriminatory power. FPG was identified as the most important predictor for I-IGT, followed by HDL, TC, HBP, BMI, and WHR. Individuals with FPG levels higher than 5.1 mmol/L were more likely to have I-IGT; the performance metrics for this cut-off value were: 89.35% accuracy, 89.79% recall, 85.22% F1 score, 81.09% PPV, 94.38% NPV, and 0.95 AUC. Machine learning models based on demographic and clinical characteristics offer a cost-effective method for predicting I-IGT in Chinese Han men aged over 50, without the need for an OGTT. These models could complement existing early diagnostic strategies, thereby enhancing the early detection and prevention of diabetes. Additionally, FPG alone could serve as an efficient screening tool for the early identification of I-IGT in clinical settings.

  • Research Article
  • 10.1038/s41598-026-51299-z
Supervised machine learning algorithms for classifications of gender-based violence in Somalia: a comparison of oversampling techniques.
  • May 28, 2026
  • Scientific reports
  • Seyifemickael Amare Yilema + 5 more

Gender-based violence can include sexual, physical, mental, and economic harm inflicted in public or in private. This violence also has a direct psychological effect, physical and financial consequences, and it has multiple underlying reasons, such as social, economic, cultural, political, and religious aspects. By applying multiple resampling techniques, this study aims to improve the precision and accuracy of supervised machine learning classifications of gender-based violence (GBV) using the SDHS dataset. The class imbalance between GBV-positive and GBV-negative instances makes it very challenging to produce reliable classification machine learning models. To address this issue, oversampling machine learning approaches, including synthetic minority over-sampling technique (SMOTE), adaptive synthetic (ADASYN), and random over-sampling (ROS), were employed to classify the GBV data in Somalia. The logistic regression (LR), decision tree (CART), random forest (RF), naïve Bayes (NB), k-nearest Neighbors (KNN), and support vector machine (SVM) methods were trained and evaluated. In addition, oversampling techniques were employed for improving the imbalanced datasets. Receiver operating characteristic curve (ROC) and the area under the curve (AUC) were used to assess each machine learning classifier and to compare performance on the original GBV dataset. Among the resampling techniques, SMOTE (RF = 0.992, CART = 0.969, and KNN = 0.957) outperformed ADASYN (RF = 0.912, CART = 0.910, and KNN = 0.876) and ROS (RF = 0.920, CART = 0.919, and KNN = 0.880) across almost all evaluation metrics. The classifiers that performed the best were random forest (RF) and classification and regression trees (CART), then k-nearest Neighbors. After resampling the imbalanced dataset, we may therefore conclude that the random forest (AUC = 0.972), CART (AUC = 0.969) and KNN (AUC = 0.957) machine learning classifiers are better at accurately classifying the k-nearest Neighbors dataset. In addition, compared to the other oversampling techniques, SMOTE was used to the machine learning classifiers to balance the imbalanced class distributions in favour of the minority class. In addition, SMOTE with the Mathews correlation coefficient (MCC) outperformed ADASYN and ROS resampling techniques. The MCC values for SMOTE reached their highest values (RF = 0.86, CART = 0.85, and KNN = 0.80), indicating strong overall predictive reliability of the machine learning models. Therefore, the findings of this analysis will assist government and non-government organizations in making policy decisions to GBV risks.

  • Research Article
  • Cite Count Icon 95
  • 10.1007/s42979-020-00296-8
Applications of Machine Learning Techniques to Predict Diagnostic Breast Cancer
  • Aug 14, 2020
  • SN Computer Science
  • Vikas Chaurasia + 1 more

This article compares six machine learning (ML) algorithms: Classification and Regression Tree (CART), Support Vector Machine (SVM), Naive Bayes (NB), K-Nearest Neighbors (KNN), Linear Regression (LR) and Multilayer Perceptron (MLP) on the Wisconsin Diagnostic Breast Cancer (WDBC) dataset by estimating their classification test accuracy, standardized data accuracy and runtime analysis. The main objective of this study is to improve the accuracy of prediction using a new statistical method of feature selection. The data set has 32 features, which are reduced using statistical techniques (mode), and the same measurements as above are applied for comparative studies. In the reduced attribute data subset (12 features), we applied 6 integrated models AdaBoost (AB), Gradient Boosting Classifier (GBC), Random Forest (RF), Extra Tree (ET) Bagging and Extra Gradient Boost (XGB), to minimize the probability of misclassification based on any single induced model. We also apply the stacking classifier (Voting Classifier) ​​to basic learners: Logistic Regression (LR), Decision Tree (DT), Support-vector clustering (SVC), K-Nearest Neighbors (KNN), Random Forest (RF) and Naive Bays (NB) to find out the accuracy obtained by voting classifier (Meta level). To implement the ML algorithm, the data set is divided in the following manner: 80% is used in the training phase and 20% is used in the test phase. To adjust the classifier, manually assigned hyper-parameters are used. At different stages of classification, all ML algorithms perform best, with test accuracy exceeding 90% especially when it is applied to a data subset.

Save Icon
Up Arrow
Open/Close
Notes

Save Important notes in documents

Highlight text to save as a note, or write notes directly

You can also access these Documents in Paperpal, our AI writing tool

Powered by our AI Writing Assistant