Lightweight Machine Learning Models For Large-Scale Heart Disease Prediction In Resource-Constrained Healthcare Environments
Lightweight Machine Learning Models For Large-Scale Heart Disease Prediction In Resource-Constrained Healthcare Environments
- Research Article
- 10.58888/2957-3912-20220105
- Jan 1, 2022
- Journal of Anesthesia and Translational Medicine
Development of A Machine Learning Model for Predicting Unanticipated Difficult Tracheal Intubation
- Research Article
3
- 10.3760/cma.j.cn112137-20200820-02425
- May 11, 2021
- Zhonghua yi xue za zhi
Objective: To explore the value of machine learning models in preoperative prediction of microvascular invasion (MVI) in hepatocellular carcinoma (HCC) based on dual-phase contrast-enhanced CT radiomics features. Methods: The data of 148 patients [106 males and 42 females, with an average age of (58±11) years] with HCC confirmed by pathology in the First Affiliated Hospital of Soochow University from January 2015 to May 2020 were retrospectively analyzed, including 88 cases of positive MVI and 60 cases of negative MVI. According to the ratio of 7∶3, the patients were randomly divided into the training and validation sets, respectively. The three-dimensional (3D) radiomics features of HCC in arterial phase (AP) and portal venous phase (PP) were extracted by MaZda software, and the optimal feature subset was obtained by combining three feature selection methods (FPM method) and Lasso regression. Then, six machine learning methods were used to build the prediction models. Receiver operating characteristic (ROC) curves were drawn to evaluate the prediction ability of the aforementioned models, and the area under the curve (AUC), accuracy, sensitivity and specificity were calculated. Results: Radiomics features of HCC in AP and PP were extracted by MaZda software, with 239 in each phase. There were 7 optimal features in AP and 14 optimal features in PP selected by FPM method and Lasso regression, respectively. The AUCs of decision tree, extreme gradient boosting, random forest, support vector machine (SVM), generalized linear model, and neural network based on the 7 optimal features in AP in the validation set were 0.736, 0.910, 0.913, 0.915, 0.897, 0.648, respectively. The SVM had the highest AUC in the validation set, with the accuracy, sensitivity and specificity of 95.35%, 95.83% and 94.74%, respectively. Likewise, the AUCs of machine learning models in prediction of MVI in HCC based on the 14 optimal features in PP in the validation set were 0.873, 0.876, 0.913, 0.859, 0.877, 0.834, respectively, and there were no significant differences (all P>0.05). The random forest had the highest AUC in the validation set, with the accuracy, sensitivity and specificity of 90.70%, 87.50% and 94.74%, respectively. Conclusion: Machine learning models based on dual-phase enhanced CT radiomics features can be used in preoperative prediction of MVI in HCC, particularly the SVM and random forest models have high prediction efficiency.
- Research Article
- 10.1051/e3sconf/202568000010
- Jan 1, 2025
- E3S Web of Conferences
This research evaluates machine learning models in predicting solar radiation, crucial for designing photovoltaic systems. Accuracy in solar forecasting is key to mitigating climate change and meeting energy demand. Advanced machine learning techniques were applied, surpassing traditional models in precision and efficiency, including SARIMA, Random Forests, SVM, ANN, and LSTM, assessed with metrics such as accuracy, sensitivity, precision, NME, R2, and execution time. After normalization, the SVM model achieved the highest overall score of 5.86. A photovoltaic system was sized using an SVM model with solar radiation data (2017-2020). Predictions calculated an average daily consumption of 4.89 kWh, a total daily energy of 109.88 kWh, and a solar panel area of 4.42 m2. The system’s peak power is 0.86 kWp, and the inverter power with a safety margin is 1.04 kW.
- Research Article
- 10.47191/etj/v9i10.17
- Oct 30, 2024
- Engineering and Technology Journal
This study investigates the effectiveness of machine learning models in predicting audit opinions using a dataset from the FiinPro-X platform, comprising 9,783 audited consolidated financial statements from public companies listed on Vietnamese stock exchanges from 2016 to 2023. The dataset spans various industries, excluding banks and financial institutions, and focuses on identifying key financial, non-financial, and qualitative variables that influence audit opinions. Six supervised learning algorithms were applied—Logistic Regression, K-Nearest Neighbors (KNN), Decision Trees, Random Forests, Support Vector Machines (SVM), and Naive Bayes—evaluated based on their ability to predict both fully acceptable (unqualified) and non-fully acceptable audit opinions. All data processing and model training were implemented in a Python environment. The Random Forest model demonstrated the best overall performance, achieving an accuracy of 0.868 and an AUC-ROC of 0.87, though its F1 score for predicting non-fully acceptable audit opinions was lower (0.585). This suggests that while machine learning models can improve prediction accuracy, challenges remain in handling imbalanced data and non-linear relationships among input variables. The study also reduced the number of features by 30%, improving the models’ performance. Future research should further refine data and feature construction processes to ensure comparability and practical applicability.
- Research Article
12
- 10.1080/10408363.2025.2497843
- May 3, 2025
- Critical Reviews in Clinical Laboratory Sciences
Despite advancements in medical care, acute kidney injury (AKI) remains a major contributor to adverse patient outcomes and presents a significant challenge due to its associated morbidity, mortality, and financial cost. Machine learning (ML) is increasingly being recognized for its potential to transform AKI care by enabling early prediction, detection, and facilitating an individualized approach to patient management. This scoping review aims to provide a comprehensive analysis of externally validated ML models for the prediction, detection, and management of AKI. We systematically searched for relevant literature from inception to 15 February 2024, using four databases—MEDLINE, EMBASE, Web of Science, and Scopus. We focused solely on models that had undergone external validation, employed Kidney Disease Improving Global Outcomes (KDIGO) definitions for AKI, and utilized ML models (excluding logistic regression models). A total of 44 studies encompassing 161 ML models for AKI prediction, severity assessment, and outcomes in both adult and pediatric populations were included in the review. These studies encompassed 4,153,424 patient admissions, with 1,209,659 in the development and internal validation cohorts and 2,943,765 in the external validation cohorts. The ML models demonstrated significant variability in performance owing to differing clinical settings, populations, and predictors used. Most of the included models were developed in specialized patient populations, such as those in intensive care units, post-surgical settings, and specific disease states (e.g. congestive heart failure, traumatic brain injury, etc.). Moreover, only a few models incorporated dynamic predictors of AKI which are crucial for improving clinical utility in rapidly evolving clinical conditions like AKI. The variable performance of these models when applied to external validation cohorts highlights the challenges of reproducibility and generalizability in implementing ML models in AKI care. Despite acceptable performance metrics, none of the models assessed in this review underwent validation or implementation in real-world clinical workflows. These findings underscore the need for standardized performance metrics and validation protocols to enhance the generalizability and clinical applicability of these models. Future efforts should focus on enhancing model adaptability by incorporating dynamic predictors and unstructured data and by ensuring that models are developed in diverse patient populations. Moreover, collaboration between clinicians and data scientists is critical to ensure the development of models that are clinically relevant, fair, and tailored to real-world healthcare environments.
- Research Article
- 10.17798/bitlisfen.1646155
- Sep 30, 2025
- Bitlis Eren Üniversitesi Fen Bilimleri Dergisi
This study investigates the effectiveness of advanced machine learning models in predicting IQ levels using a diverse set of socioeconomic and health indicators from global databases such as WHO, the World Bank, and United Nations organizations. The research employs various algorithms, including Linear Regression, Random Forest, Gradient Boosting, Support Vector Machines, Ridge and Lasso Regressions, XGBoost, LightGBM, and a Stacking Regressor to capture both linear and non-linear relationships. Evaluation metrics such as Mean Absolute Error (MAE), Root Mean Square Error (RMSE), and R-squared (R²) reveal that LightGBM and Stacking Regressor models excel in accuracy and generalization. The study highlights the trade-off between model interpretability and predictive power, emphasizing that simpler models offer greater transparency. In contrast, more complex models successfully capture intricate interactions among education, health, and economic factors. The findings provide valuable insights for policymakers and researchers, suggesting that machine learning approaches can significantly enhance understanding of the determinants of IQ and aid in developing targeted strategies in education and social policy.
- Research Article
- 10.5005/jp-journals-10071-25157
- Apr 1, 2026
- Indian Journal of Critical Care Medicine : Peer-reviewed, Official Publication of Indian Society of Critical Care Medicine
A<sc>bstract</sc>Background and aimsPrecise forecasting of mortality in intensive care units (ICUs) is essential for enhancing patient management and resource distribution. Traditional scoring methods like the Acute Physiology and Chronic Health Evaluation II (APACHE II) and Sequential Organ Failure Assessment (SOFA) are prevalent yet may inadequately encompass the intricacies of critical illness. The aim of this work was to create and internally test machine learning models for predicting mortality in the ICU, utilizing routinely gathered electronic health record data.Patients and methodsThis retrospective cohort analysis encompassed 5,553 adult ICU hospitalizations from September 2021 to December 2023. Patients were randomly allocated into development (80%) and test (20%) cohorts. Three predictive models—Logistic Regression, Random Forest, and Extreme Gradient Boosting (XGBoost)—were developed utilizing demographic information, APACHE II and SOFA values, comorbidities, and ventilatory support status. The evaluation of model performance was conducted utilizing the area under the receiver operating characteristic curve (AUROC), precision, recall, and F1 score.ResultsAmong the 5,553 analyzed ICU admissions, 1,869 patients succumbed during their ICU stay, yielding an ICU mortality rate of 33.6% (95% CI: 32.4–34.9). The Random Forest model had the superior discriminative capability for mortality prediction, achieving an AUROC of 0.842, followed by XGBoost with an AUROC of 0.835, and Logistic Regression with an AUROC of 0.833. Although Logistic Regression demonstrated slightly superior overall accuracy, ensemble models more effectively identified non-linear correlations among predictors. Acute Physiology and Chronic Health Evaluation II and SOFA values proved to be the most significant predictors in all models. Temporal validation and sensitivity analysis produced consistent outcomes, demonstrating the resilience of model performance.ConclusionMachine learning models exhibited strong efficacy in predicting ICU mortality, with Random Forest and XGBoost displaying slight advantages over Logistic Regression. The incorporation of machine learning algorithms into existing ICU scoring systems may improve risk categorization and facilitate clinical decision-making.How to cite this articleDash A, Majhi K, Ghosh N, Jena P, Choudhury S, Anushapreethi S, et al. Development and Validation of a Multivariable Machine Learning Model for Mortality Prediction among Intensive Care Unit Patients. Indian J Crit Care Med 2026;30(4):282–290.
- Supplementary Content
1
- 10.1016/j.ijcrp.2026.200619
- Mar 6, 2026
- International Journal of Cardiology. Cardiovascular Risk and Prevention
Machine learning and Regression–Based models for prediction of postoperative atrial fibrillation following coronary artery bypass grafting: A systematic review and meta-analysis
- Research Article
1
- 10.52783/jisem.v10i31s.5137
- Apr 2, 2025
- Journal of Information Systems Engineering and Management
A crucial aspect of financial markets is volatility forecasting, which enables analysts, investors, and policymakers to assess risk, optimize portfolios, and develop trading plans. This research investigates which machine learning (ML) model is best fit for predicting sectoral volatility. Comparing models like Random Forest, Gradient Boosting, Neural Networks, and Support Vector Regression, we apply historical data from 2012 to 2024 from eleven different sectors or industries including FMCG, Energy, Financial Services, auto etc. The research is unique in terms of holistic view of all Indian sectoral volatility indices and relationship between them. Moreover, no research has identified or compared the best ML models for predicting sectoral volatility. With the greatest R square value of 0.998 and the lowest Mean Absolute Error (MAE) of 0.765, our results demonstrate that Random Forest outperforms the other models. Significant sectoral connections and volatility patterns are also shown in the analysis, especially during the COVID-19 pandemic like sectors more susceptible to economic shocks, such as financial services and fast-moving consumer goods. Heatmaps and time-series visualization techniques demonstrate that sectoral interdependence and volatility clustering. These results highlight the potential of machine learning techniques to improve risk management and volatility pre-diction and provide insightful data for financial analysts, investors, and policymakers. The study recommends applying Random Forest ML models for volatility prediction and investment decision-making, considering limitations such as dataset bias and challenges in predicting extreme market conditions.
- Research Article
- 10.12122/j.issn.1673-4254.2026.01.15
- Jan 20, 2026
- Nan fang yi ke da xue xue bao = Journal of Southern Medical University
To improve the accuracy of machine learning models for preoperative prediction of high-intensity focused ultrasound (HIFU) ablation efficacy for uterine fibroids by correcting class imbalance in small sample datasets using undersampling methods. Clinical and imaging data were collected from 140 patients with uterine fibroids undergoing HIFU treatment at Foshan Women and Children Hospital, including 104 with high ablation rates and 36 with low ablation rates. Radiomic features were extracted from MRI T2-weighted images (T2WI) of the patients, and machine learning models were constructed to predict HIFU treatment outcomes. Four machine learning algorithms, including k-Nearest Neighbors (KNN), Random Forest (RF), Support Vector Machine (SVM), and Multilayer Perceptron (MLP), were coupled with 7 undersampling methods, namely Random Undersampling (RUS), Repeated Edited Nearest Neighbors (RENN), All k-Nearest Neighbors (AllKNN), Neighborhood Cleaning Rule-3 (NM), Condensed Nearest Neighbor (CNN), Neighborhood Cleaning Rule (NCR), and Instance Hardness Threshold (IHT), for handling class imbalance in the datasets. The 28 prediction models were evaluated using 5-fold cross-validation for areas under the receiver operating characteristic curve (AUC), accuracy, recall, and specificity. The best combinations of undersampling methods and machine learning models CNN-RF, NM-SVM, CNN-KNN, and NM-MLP had AUCs of 0.772 (95% CI: 0.566-0.942), 0.797 (95% CI: 0.600-0.950), 0.822 (95% CI: 0.635-0.964), and 0.822 (95% CI: 0.632-0.960), respectively. The AUCs of the machine learning models significantly increased after coupling with undersampling methods, with the MLP model showing the most pronounced improvement. The recall rates of the 4 combined models also improved significantly (by 0.389 for CNN-RF, 0.836 for NM-SVM, 0.532 for CNN-KNN, and 0.372 for NM-MLP). The use of undersampling methods can effectively correct class imbalance in small sample datasets to improve the accuracy of machine learning models for predicting the efficacy of HIFU ablation for uterine fibroids.
- Research Article
5
- 10.1186/s12302-024-00841-9
- Jan 12, 2024
- Environmental Sciences Europe
Monitoring water resources requires accurate predictions of rainfall data. Our study introduces a novel deep learning model named the deep residual shrinkage network (DRSN)—temporal convolutional network (TCN) to remove redundant features and extract temporal features from rainfall data. The TCN model extracts temporal features, and the DRSN enhances the quality of the extracted features. Then, the DRSN–TCN is coupled with a random forest (RF) model to model rainfall data. Since the RF model may be unable to classify and predict complex patterns and data, our study develops the RF model to model outputs with high accuracy. Since the DRSN–TCN model uses advanced operators to extract temporal features and remove irrelevant features, it can improve the performance of the RF model for predicting rainfall. We use a new optimizer named the Gaussian mutation (GM)–orca predation algorithm (OPA) to set the DRSN–TCN–RF (DTR) parameters and determine the best input scenario. This paper introduces a new machine learning model for rainfall prediction, improves the accuracy of the original TCN, and develops a new optimization method for input selection. The models used the lagged rainfall data to predict monthly data. GM–OPA improved the accuracy of the orca predation algorithm (OPA) for feature selection. The GM–OPA reduced the root mean square error (RMSE) values of OPA and particle swarm optimization (PSO) by 1.4%–3.4% and 6.14–9.54%, respectively. The GM–OPA can simplify the modeling process because it can determine the most important input parameters. Moreover, the GM–OPA can automatically determine the optimal input scenario. The DTR reduced the testing mean absolute error values of the TCN–RAF, DRSN–TCN, TCN, and RAF models by 5.3%, 21%, 40%, and 46%, respectively. Our study indicates that the proposed model is a reliable model for rainfall prediction.
- Research Article
3
- 10.1186/s12911-025-03142-0
- Aug 8, 2025
- BMC Medical Informatics and Decision Making
BackgroundAnti-programmed cell death protein 1 (PD-1)/programmed cell death ligand 1 (PD-L1) immunotherapy has revolutionized cancer treatment. However, it can cause immune-related adverse events, including acute kidney injury (AKI). Such adverse events can interrupt treatment, affecting patient outcomes. Early prediction of AKI is essential for improved prognosis and personalized therapeutic strategies. Previous research has been constrained by significant limitations, underscoring the necessity for AKI risk prediction models for patients treated with PD-1/PD-L1 inhibitors. This study aimed to develop and validate an interpretable machine learning (ML) model for early AKI prediction in patients undergoing PD-1/PD-L1 inhibitor therapy using a retrospective cohort design.MethodsThis study collected data from patients treated with PD-1/PD-L1 inhibitors at Zhejiang Provincial People’s Hospital between January 2018 and January 2024. Nine ML models were evaluated. SHapley Additive exPlanations (SHAP) were employed to rank feature importance and interpret the final model. Additionally, a web-based calculator based on the model was developed.ResultsAmong the nine ML models evaluated, the Grandient Boosting Machine (GBM) model achieved the best predictive performance. In the validation set, the GBM model achieved an AUC of 0.850 (95%CI: 0.830–0.870). In the test set, the AUC was 0.795(95% CI: 0.747–0.844), demonstrating accurate AKI risk prediction. Calibration curves demonstrated a strong concordance between predicted and observed risk probabilities. An interpretable final GBM model with 13 features was developed after feature reduction based on feature importance ranking. A web-based calculator accessible at https://predicatingaki.shinyapps.io/PDmodel/ has been developed to assist clinicians in AKI risk assessment.ConclusionThis study developed and validated an interpretable ML model using a large dataset to predict AKI risk in patients receiving PD-1/PD-L1 inhibitor therapy. This model can assist clinicians in the early identification of high-risk patients, facilitating personalized treatment plans.Trial registrationThe study was conducted following the Declaration of Helsinki and was approved by the Ethics Committee of Zhejiang Provincial People’s Hospital (Approval No. KT2024116) in 3 Jan. 2025. As it was a retrospective study with anonymized data, informed consent was waived.Supplementary InformationThe online version contains supplementary material available at 10.1186/s12911-025-03142-0.
- Research Article
56
- 10.1016/j.eclinm.2024.102913
- Oct 30, 2024
- eClinicalMedicine
Development and validation of an interpretable machine learning model for predicting the risk of distant metastasis in papillary thyroid cancer: a multicenter study
- Research Article
1
- 10.30591/jpit.v9i2.7051
- Aug 30, 2024
- Jurnal Informatika: Jurnal Pengembangan IT
Solar panels have become a popular source of renewable energy due to their sustainability and environmental friendliness. Accurate predictions of solar panel output are crucial for various applications, such as energy system optimization, power grid management, and economic planning. Many important factors pose challenges in predicting the output of solar panels, such as weather conditions that can change at any time, geographical factors, data quality, and the duration of data collection. Machine learning (ML) models show promising performance in this prediction; there are many types of machine learning models, some are single models and others are hybrid models. Optimization algorithms are used to optimize parameters and improve the prediction accuracy of machine learning models. This research reviews fifteen journals that have been filtered to obtain those discussing optimization algorithms in the predictive models of solar panel output power. This journal will examine the optimization algorithms used in machine learning models for predicting solar panel output power, discussing various types of optimization algorithms, their application in machine learning models, the prediction results from these models, the input data used, and the data collection locations that significantly influence the prediction outcomes. From the results of this research, it does not conclude which machine learning model is the best, due to the many factors that influence it. However, this research is expected to provide references on the application of machine learning models in predicting the output power of solar panels, thereby encouraging the use of renewable energy sources.
- Research Article
12
- 10.1002/cam4.6901
- Jan 1, 2024
- Cancer Medicine
ObjectivesDevelopment and validation of a computed tomography urography (CTU)‐based machine learning (ML) model for prediction of preoperative pathology grade of upper urinary tract urothelial carcinoma (UTUC).MethodsA total of 140 patients with UTUC who underwent CTU examination from January 2017 to August 2023 were retrospectively enrolled. Tumor lesions on the unenhanced, medullary, and excretory periods of CTU were used to extract Features, respectively. Feature selection was screened by the Pearson and Spearman correlation analysis, least absolute shrinkage and selection operator algorithm, random forest (RF), support vector machine (SVM), and eXtreme Gradient Boosting (XGBoost). The logistic regression (LR) was used to screen for independent influencing factors of clinical baseline characteristics. Machine learning models based on different feature datasets were constructed and validated using algorithms such as LR, RF, SVM, and XGBoost. By computing the selected features, a radiomics score was generated, and a diverse feature dataset was constructed. Based on the training set, 16 ML models were created, and their performance was evaluated using the validation set for metrics including sensitivity, specificity, accuracy, area under the receiver operating characteristic curve (AUC), and others.ResultsThe training set consisted of 98 patients (mean age: 64.5 ± 10.5 years; 30 males), whereas the validation set consisted of 42 patients (mean age: 65.3 ± 9.78 years; 17 males). Hydronephrosis was the best independent influence factor (p < 0.05). The RF model had the best performance in predicting high‐grade UTUC, with AUC of 0.914 (95% Confidence Interval [95%CI] 0.852–0.977) and 0.903 (95%CI 0.809–0.997) in the training set and validation set, and accuracy of 0.878 and 0.857, respectively.ConclusionsAn ML model based on the RF algorithm exhibits excellent predictive performance, offering a non‐invasive approach for predicting preoperative high‐grade UTUC.