Optimal ratio for data splitting
Abstract It is common to split a dataset into training and testing sets before fitting a statistical or machine learning model. However, there is no clear guidance on how much data should be used for training and testing. In this article, we show that the optimal training/testing splitting ratio is , where is the number of parameters in a linear regression model that explains the data well.
- Research Article
168
- 10.1371/journal.pgen.1004754
- Nov 13, 2014
- PLoS Genetics
Compared to univariate analysis of genome-wide association (GWA) studies, machine learning-based models have been shown to provide improved means of learning such multilocus panels of genetic variants and their interactions that are most predictive of complex phenotypic traits. Many applications of predictive modeling rely on effective variable selection, often implemented through model regularization, which penalizes the model complexity and enables predictions in individuals outside of the training dataset. However, the different regularization approaches may also lead to considerable differences, especially in the number of genetic variants needed for maximal predictive accuracy, as illustrated here in examples from both disease classification and quantitative trait prediction. We also highlight the potential pitfalls of the regularized machine learning models, related to issues such as model overfitting to the training data, which may lead to overoptimistic prediction results, as well as identifiability of the predictive variants, which is important in many medical applications. While genetic risk prediction for human diseases is used as a motivating use case, we argue that these models are also widely applicable in nonhuman applications, such as animal and plant breeding, where accurate genotype-to-phenotype modeling is needed. Finally, we discuss some key future advances, open questions and challenges in this developing field, when moving toward low-frequency variants and cross-phenotype interactions.
- Research Article
48
- 10.3389/fcvm.2022.812276
- Apr 6, 2022
- Frontiers in Cardiovascular Medicine
ObjectiveTo compare the performance, clinical feasibility, and reliability of statistical and machine learning (ML) models in predicting heart failure (HF) events.BackgroundAlthough ML models have been proposed to revolutionize medicine, their promise in predicting HF events has not been investigated in detail.MethodsA systematic search was performed on Medline, Web of Science, and IEEE Xplore for studies published between January 1, 2011 to July 14, 2021 that developed or validated at least one statistical or ML model that could predict all-cause mortality or all-cause readmission of HF patients. Prediction Model Risk of Bias Assessment Tool was used to assess the risk of bias, and random effect model was used to evaluate the pooled c-statistics of included models.ResultTwo-hundred and two statistical model studies and 78 ML model studies were included from the retrieved papers. The pooled c-index of statistical models in predicting all-cause mortality, ML models in predicting all-cause mortality, statistical models in predicting all-cause readmission, ML models in predicting all-cause readmission were 0.733 (95% confidence interval 0.724–0.742), 0.777 (0.752–0.803), 0.678 (0.651–0.706), and 0.660 (0.633–0.686), respectively, indicating that ML models did not show consistent superiority compared to statistical models. The head-to-head comparison revealed similar results. Meanwhile, the immoderate use of predictors limited the feasibility of ML models. The risk of bias analysis indicated that ML models' technical pitfalls were more serious than statistical models'. Furthermore, the efficacy of ML models among different HF subgroups is still unclear.ConclusionsML models did not achieve a significant advantage in predicting events, and their clinical feasibility and reliability were worse.
- Conference Article
- 10.3390/ecws-4-06441
- Nov 12, 2019
Pipe failures in Water Distribution Networks (WDN) may cause economic, environmental and social costs. The application of statistical and Machine Learning (ML) models play a critical role in planning and decision support processes for WDN management. Failure models can provide valuable information for prioritizing the system rehabilitation even in data scarcity scenarios (such as developing countries). This study compares several statistical and ML pipe failure models thus providing useful information to practitioners to select a suitable model according to their needs. Three statistical models (i.e. Linear, Poisson and Evolutionary Polynomial Regressions) were used for pipe failures prediction based on diameter, age of pipes and length as explanatory variables. The K-means clustering approach was applied to improve the performance of the statistical models. The performance indicators used were the coefficient of determination (R2) and the root mean square error (RMSE). ML approaches - namely Gradient Boosted Tree (GBT), Bayes, Support Vector Machine and Artificial Neuronal Networks (ANNs) - were compared in predicting individual pipe failure rates. The pipe's attributes, environmental and operational variables were included as input variables. Their performance was evaluated using confusion matrices and receiver operating characteristic curves. The proposed approach was applied to a WDN in Bogotá (Colombia). The results showed that the cluster-based prediction model reduces the prediction error of pipe failures. All the models demonstrated acceptable results in terms of their performance (R2 between 0.695-0.927 and RMSE between 45-22 for the test sample). Regarding ML models, all methods but the ANNs show acceptable performance. The GBT approach has the best performing classifier (79.41% correct predictions in the test sample). This model was used to calculate the failure rate of individual pipes for rehabilitation planning. Furthermore, a sensitivity analysis of the GBT model to the input variables was performed to provide information on its generalization capability.
- Research Article
18
- 10.1186/s12911-023-02166-8
- Apr 21, 2023
- BMC Medical Informatics and Decision Making
ObjectivesThis research was designed to compare the ability of different machine learning (ML) models and nomogram to predict distant metastasis in male breast cancer (MBC) patients and to interpret the optimal ML model by SHapley Additive exPlanations (SHAP) framework.MethodsFour powerful ML models were developed using data from male breast cancer (MBC) patients in the SEER database between 2010 and 2015 and MBC patients from our hospital between 2010 and 2020. The area under curve (AUC) and Brier score were used to assess the capacity of different models. The Delong test was applied to compare the performance of the models. Univariable and multivariable analysis were conducted using logistic regression.ResultsOf 2351 patients were analyzed; 168 (7.1%) had distant metastasis (M1); 117 (5.0%) had bone metastasis, and 71 (3.0%) had lung metastasis. The median age at diagnosis is 68.0 years old. Most patients did not receive radiotherapy (1723, 73.3%) or chemotherapy (1447, 61.5%). The XGB model was the best ML model for predicting M1 in MBC patients. It showed the largest AUC value in the tenfold cross validation (AUC:0.884; SD:0.02), training (AUC:0.907; 95% CI: 0.899—0.917), testing (AUC:0.827; 95% CI: 0.802—0.857) and external validation (AUC:0.754; 95% CI: 0.739—0.771) sets. It also showed powerful ability in the prediction of bone metastasis (AUC: 0.880, 95% CI: 0.856—0.903 in the training set; AUC: 0.823, 95% CI:0.790—0.848 in the test set; AUC: 0.747, 95% CI: 0.727—0.764 in the external validation set) and lung metastasis (AUC: 0.906, 95% CI: 0.877—0.928 in training set; AUC: 0.859, 95% CI: 0.816—0.891 in the test set; AUC: 0.756, 95% CI: 0.732—0.777 in the external validation set). The AUC value of the XGB model was larger than that of nomogram in the training (0.907 vs 0.802) and external validation (0.754 vs 0.706) sets.ConclusionsThe XGB model is a better predictor of distant metastasis among MBC patients than other ML models and nomogram; furthermore, the XGB model is a powerful model for predicting bone and lung metastasis. Combining with SHAP values, it could help doctors intuitively understand the impact of each variable on outcome.
- Research Article
142
- 10.1007/s10661-019-7330-6
- Mar 5, 2019
- Environmental Monitoring and Assessment
Spatio-temporal land-use change modeling, simulation, and prediction have become one of the critical issues in the last three decades due to uncertainty, structure, flexibility, accuracy, the ability for improvement, and the capability for integration of available models. Therefore, many types of models such as dynamic, statistical, and machine learning (ML) models have been used in the geographic information system (GIS) environment to fulfill the high-performance requirements of land-use modeling. This paper provides a literature review on models for modeling, simulating, and predicting land-use change to determine the best approach that can realistically simulate land-use changes. Therefore, the general characteristics of conventional and ML models for land-use change are described, and the different techniques used in the design of these models are classified. The strengths and weaknesses of the various dynamic, statistical, and ML models are determined according to the analysis and discussion of the characteristics of these models. The results of the review confirm that ML models are the most powerful models for simulating land-use change because they can include all driving forces of land-use change in the simulation process and simulate linear and non-linear phenomena, which dynamic models and statistical models are unable to do. However, ML models also have limitations. For instance, some ML models are complex, the simulation rules cannot be changed, and it is difficult to understand how ML models work in a system. However, this can be solved via the use of programming languages such as Python, which in turn improve the simulation capabilities of the ML models.
- Research Article
3
- 10.1167/tvst.13.8.12
- Aug 8, 2024
- Translational vision science & technology
Compare the use of optic disc and macular optical coherence tomography measurements to predict glaucomatous visual field (VF) worsening. Machine learning and statistical models were trained on 924 eyes (924 patients) with circumpapillary retinal nerve fiber layer (cp-RNFL) or ganglion cell inner plexiform layer (GC-IPL) thickness measurements. The probability of 24-2 VF worsening was predicted using both trend-based and event-based progression definitions of VF worsening. Additionally, the cp-RNFL and GC-IPL predictions were combined to produce a combined prediction. A held-out test set of 617 eyes was used to calculate the area under the curve (AUC) to compare cp-RNFL, GC-IPL, and combined predictions. The AUCs for cp-RNFL, GC-IPL, and combined predictions with the statistical and machine learning models were 0.72, 0.69, 0.73, and 0.78, 0.75, 0.81, respectively, when using trend-based analysis as ground truth. The differences in performance between the cp-RNFL, GC-IPL, and combined predictions were not statistically significant. AUCs were highest in glaucoma suspects using cp-RNFL predictions and highest in moderate/advanced glaucoma using GC-IPL predictions. The AUCs for the statistical and machine learning models were 0.63, 0.68, 0.69, and 0.72, 0.69, 0.73, respectively, when using event-based analysis. AUCs decreased with increasing disease severity for all predictions. cp-RNFL and GC-IPL similarly predicted VF worsening overall, but cp-RNFL performed best in early glaucoma stages and GC-IPL in later stages. Combining both did not enhance detection significantly. cp-RNFL best predicted trend-based 24-2 VF progression in early-stage disease, while GC-IPL best predicted progression in late-stage disease. Combining both features led to minimal improvement in predicting progression.
- Research Article
- 10.7759/cureus.91318
- Aug 30, 2025
- Cureus
Background and objectivesIn the past twenty years, several large-scale coronavirus outbreaks have caused heavy loss of life and serious economic damage worldwide. Current global surveillance suggests that similar epidemics may occur again, making timely and accurate forecasting an urgent priority. Yet, many existing prediction methods, mainly based on traditional statistical or machine learning techniques, still struggle to deliver both speed and precision. This study explores a generative artificial intelligence-driven approach aimed at narrowing these gaps.MethodsNine models (three statistical models, three machine learning models, and three generative artificial intelligence models) were compared using weekly COVID-19 case and death data from the United States (US), the United Kingdom (UK), Germany (GE), and Russia (RU) from March 15, 2020, to April 15, 2023. The statistical models used are simple moving average (SMA), simple exponential smoothing (SES), and the Holt linear trend model (Holt). The machine learning models used are k-nearest neighbor regression (KNN), regression tree (RTree), and multilayer perceptron (MLP). The generative AI models used are ChatGPT, DeepSeek (DS), and Kimi. A custom MATLAB program was used to solve the statistical and machine learning models, and the zero-inference forecasting method was used to solve the generative AI model. According to the stepwise prediction theory, error metrics for one-, two-, and three-step forecasts were calculated: mean absolute percentage error (MAPE), mean absolute error (MAE), and root mean square error (RMSE). The forecasting performance of each model was compared by comparing the one-, two-, and three-step predicting error metrics.ResultsIn our analysis, generative AI models consistently delivered the most accurate forecasts. Kimi, in particular, recorded the smallest errors for death predictions and among the lowest for new cases, while DS and ChatGPT also performed well, clearly surpassing the statistical and machine learning approaches in short-term COVID-19 forecasting.ConclusionThe results of this study demonstrate that generative AI models demonstrate superior predictive accuracy and robustness in epidemic forecasting compared to traditional statistical and machine learning models. This research is innovative in its application of generative AI technology to public health decision-making, demonstrating its robust epidemic forecasting capabilities. Given these proven advantages, public health authorities can integrate generative AI technology into major infectious disease surveillance systems, promote public health data sharing mechanisms, and incorporate generative AI into epidemic intervention and resource allocation. The implementation of these measures will enable governments and regulatory agencies worldwide to use generative AI to enhance early warning capabilities and improve their response to future infectious disease epidemics.
- Research Article
2
- 10.1002/cnm.70029
- Mar 1, 2025
- International journal for numerical methods in biomedical engineering
The complex mechanical environment of peripheral arteries makes stents with poor torsional performance more prone to fracture, and stent fracture is considered a precursor to in-stent restenosis (ISR). Therefore, studying the torsional performance of stents is crucial. However, while the finite element method (FEM) can accurately simulate the torsional behavior of stents, its time-consuming nature makes it difficult to meet the rapid design requirements for individualized stents. Thus, integrating efficient machine learning (ML) models into the stent design process may be a viable approach. In this study, a machine learning-based rapid prediction method was established to achieve the rapid prediction of torsional performance of personalized peripheral artery stents. A dataset containing 200 different stent designs was generated using Latin Hypercube Sampling (LHS) and FEM. The dataset was divided into a training set (160 samples) and a test set (40 samples). Based on four input variables-the length of strut ring (LS), the width of strut (WS), the width of link (WL), and the thickness of stent (T)-the predictive performance of polynomial regression (PR), random forest regression (RFR), and support vector regression (SVR) for the twist metric (TM) was compared. To simulate the real-world application of ML models, after training and testing the ML models, the entire dataset (combining the training and test sets) was used for re-learning while keeping the control parameters unchanged. A validation set (10 samples) was generated through sampling and FEM, and the re-learned ML models were used to predict and validate their performance. By comprehensively comparing the predictive performance of the ML models on the training set, test set, and validation set, the algorithm performance ranked as follows: PR>SVR>RFR. The PR model achieved a mean absolute error (MAE) of (training set = 0.02847; test set = 0.03083; validation set = 0.04311) and a coefficient of determination (R2) of (training set = 0.95148; test set = 0.97822; validation set = 0.94397). This method can effectively shorten the design cycle of stents and meet the need for personalized stent rapid design and choice. In addition, this method can also be extended to predict other mechanical properties of the stent and can be used in stent multi-objective design optimization.
- Research Article
66
- 10.1155/2020/9628957
- Jan 20, 2020
- Journal of Advanced Transportation
Accurate prediction of traffic information (i.e., traffic flow, travel time, traffic speed, etc.) is a key component of Intelligent Transportation System (ITS). Traffic speed is an important indicator to evaluate traffic efficiency. Up to date, although a few studies have considered the periodic feature in traffic prediction, very few studies comprehensively evaluate the impact of periodic component on statistical and machine learning prediction models. This paper selects several representative statistical models and machine learning models to analyze the influence of periodic component on short-term speed prediction under different scenarios: (1) multi-horizon ahead prediction (5, 15, 30, 60 minutes ahead predictions), (2) with and without periodic component, (3) two data aggregation levels (5-minute and 15-minute), (4) peak hours and off-peak hours. Specifically, three statistical models (i.e., space time (ST) model, vector autoregressive (VAR) model, autoregressive integrated moving average (ARIMA) model) and three machine learning approaches (i.e., support vector machines (SVM) model, multi-layer perceptron (MLP) model, recurrent neural network (RNN) model) are developed and examined. Furthermore, the periodic features of the speed data are considered via a hybrid prediction method, which assumes that the data consist of two components: a periodic component and a residual component. The periodic component is described by a trigonometric regression function, and the residual component is modeled by the statistical models or the machine learning approaches. The important conclusions can be summarized as follows: (1) the multi-step ahead prediction accuracy improves when considering the periodic component of speed data for both three statistical models and three machine learning models, especially in the peak hours; (2) considering the impact of periodic component for all models, the prediction performance improvement gradually becomes larger as the time step increases; (3) under the same prediction horizon, the prediction performance of all models for 15-minute speed data is generally better than that for 5-minute speed data. Overall, the findings in this paper suggest that the proposed hybrid prediction approach is effective for both statistical and machine learning models in short-term speed prediction.
- Research Article
227
- 10.3390/info11060332
- Jun 20, 2020
- Information
Forecasting the direction and trend of stock price is an important task which helps investors to make prudent financial decisions in the stock market. Investment in the stock market has a big risk associated with it. Minimizing prediction error reduces the investment risk. Machine learning (ML) models typically perform better than statistical and econometric models. Also, ensemble ML models have been shown in the literature to be able to produce superior performance than single ML models. In this work, we compare the effectiveness of tree-based ensemble ML models (Random Forest (RF), XGBoost Classifier (XG), Bagging Classifier (BC), AdaBoost Classifier (Ada), Extra Trees Classifier (ET), and Voting Classifier (VC)) in forecasting the direction of stock price movement. Eight different stock data from three stock exchanges (NYSE, NASDAQ, and NSE) are randomly collected and used for the study. Each data set is split into training and test set. Ten-fold cross validation accuracy is used to evaluate the ML models on the training set. In addition, the ML models are evaluated on the test set using accuracy, precision, recall, F1-score, specificity, and area under receiver operating characteristics curve (AUC-ROC). Kendall W test of concordance is used to rank the performance of the tree-based ML algorithms. For the training set, the AdaBoost model performed better than the rest of the models. For the test set, accuracy, precision, F1-score, and AUC metrics generated results significant to rank the models, and the Extra Trees classifier outperformed the other models in all the rankings.
- Research Article
2
- 10.3390/medicina61020188
- Jan 22, 2025
- Medicina
Background and Objectives: Recent research has focused on exploring the relationships between various factors associated with headaches and understanding their impact on individuals’ psychological states. Utilizing statistical methods and machine learning models, these studies aim to analyze and predict these relationships to develop effective approaches for headache management and prevention. Materials and Methods: Analyzing data from 398 patients (train set = 318 and test set = 80), we investigated the influence of various features on outcomes such as depression, anxiety, and headache intensity using machine learning and linear regression. The study employed a mixed-methods approach, combining medical records, interviews, and surveys to gather comprehensive data on participants’ experiences with headaches and their associated psychological effects. Results: Machine learning models, including Random Forest (utilized for Headache Impact Test-6, Patient Health Questionnaire-9, and Generalized Anxiety Disorder-7) and Support Vector Regression (applied to Migraine Disability Assessment), revealed key features contributing to each outcome through Shapley values, while linear regression provided additional insights. Frequent analgesic medication emerged as a significant predictor of poorer life quality (Headache Impact Test-6, root mean squared error = 7.656) and increased depression (Patient Health Questionnaire-9, root mean squared error = 5.07) and anxiety (Generalized Anxiety Disorder-7, root mean squared error = 4.899) in the Random Forest model. However, interpreting the importance of features in complex models like supportive vector regression poses challenges, and determining causality between factors such as medication usage and pain severity was not feasible. Conclusions: Our study underscores the importance of considering individual characteristics in optimizing treatment strategies for headache patients.
- Research Article
1
- 10.1016/j.ijom.2025.06.020
- Feb 1, 2026
- International journal of oral and maxillofacial surgery
Machine learning approach using radiomics features to distinguish odontogenic cysts and tumours.
- Research Article
- 10.1161/circ.152.suppl_3.4342359
- Nov 4, 2025
- Circulation
Background: Despite the development of disease-modifying therapy (DMT) for patients with transthyretin amyloid cardiomyopathy (ATTR-CM), the rates of mortality and disease progression are still high. The conventional models to predict prognosis (e.g., the Columbia score, the National Amyloidosis Centre [NAC] staging) predate DMT and offer limited prediction in the era of DMT. Aim: To develop a novel prognostication model in patients with ATTR-CM using plasma proteomics. Hypothesis: Plasma proteomics improves the prediction of death and disease progression in patients with ATTR-CM beyond the conventional models. Methods: In this prospective study, we conducted plasma proteomics profiling of 7,289 proteins in patients with ATTR-CM (n=303) at enrollment. The primary outcome was all-cause death, and the secondary was disease progression–a composite of death, heart transplant, heart failure hospitalization, and oral diuretics intensification. We randomly divided the cohort into training (2/3) and test sets (1/3). In the training set, we specified proteins to predict both outcomes using the Boruta algorithm. Using the specified proteins, we developed a random forest-based machine-learning (ML) model to predict each outcome in the training set. We compared the predictive ability of the ML model with the conventional models in the test set. We performed survival analyses between high- and low-risk groups defined by the ML models in the test set, while adjusting for the Columbia score. Results: During a median follow-up of 3.9 [1 st –3 rd quartile: 2.2–5.9] years, 69 patients (23%) died and 124 (41%) had disease progression, despite 243 (80%) receiving DMT. In the training set, 17 proteins were specified to predict both outcomes. In the test set, the area under the receiver-operating-characteristic curve (AUC) of the ML model was 0.90 (95% confidence interval 0.84–0.96) for all-cause death and 0.82 (0.73–0.91) for disease progression ( Image 1 ). Each ML model outperformed the conventional models in AUC, classification, and time-dependent AUC ( Images 1, 2 ). The high-risk group in the test set specified by the ML model had a higher event rate than the low-risk group in each outcome (all-cause death, adjusted hazard ratio [aHR] 6.6 [2.5–18.0], P =0.0002; disease progression, aHR 5.7 [2.6–12.4], P <0.0001; Image 3 ). Conclusion: This study first demonstrated that plasma proteomics improves the prediction of mortality and disease progression in patients with ATTR-CM in the DMT era.
- Abstract
1
- 10.1182/blood-2024-200237
- Nov 5, 2024
- Blood
Prediction of Poor Survival after Hematopoietic Cell Transplantation in Myelofibrosis Using Machine Learning Techniques
- Research Article
- 10.1007/s11657-025-01602-8
- Oct 22, 2025
- Archives of osteoporosis
Existing osteoporosis screening tools are inaccurate and inconvenient, prompting the need for a better alternative. A machine learning tool (Gradient Boosting) with key factors (weight, age, height) outperformed OST (AUC 0.828 vs 0.781, p < 0.0001) in validation. The validated, clinically applicable tool improves osteoporosis screening accessibility and accuracy. As the first "line of defence" for osteoporosis detection, existing screening tools have low accuracy and are inconvenient to use. Therefore, this study aims to develop a machine-learning-based, clinically applicable, and interpretable osteoporosis screening tool. This study included 9405 American participants aged 50years and older (with the average age of the osteoporosis population in the training set and test set being 72 ± 9years and 73 ± 8years, respectively). The study selected 13 clinically accessible indicators as candidate predictive variables, divided the data into a training set and a test set at a ratio of 7:3, used the Lasso for feature selection, compared six statistical and machine learning models, evaluated model performance through metrics such as the Area Under the Receiver Operating Characteristic Curve (AUC), Sensitivity, specificity, F1-score, decision curve, calibration curve, and clinical impact curve, employed the SHAP (Shapley Additive exPlanations) method to enhance model interpretability, and conducted external validation based on an independent dataset from the Second Hospital of Lanzhou University. "Weight," "age," and "height" are the most critical predictive factors. Gradient Boosting Machine (GB) showed optimal results, with training and test set AUC (0.850, 0.841), sensitivity (0.757, 0.737), specificity (0.793, 0.779), and F1-score (0.336, 0.316), respectively. External validation (3500 subjects) showed that the GB-based screening tool had an AUC of 0.828, which was significantly higher than that of the traditional Osteoporosis Self-Assessment Tool (OST, AUC = 0.781) via the DeLong test (z = 10.880, p < 0.0001). A clinically applicable osteoporosis screening tool based on machine learning algorithms was developed and validated.