Evaluating Random Forest Model Performance for Cave and Sinkhole Prediction in the Cradle of Humankind, South Africa: Preliminary Analysis and Variable Importance Assessments
Abstract Surveying an area for new fossil sites is a labor-intensive and resource-draining activity that can be alleviated with the aid of machine learning models. In karst landscapes of southern Africa, Plio-Pleistocene fossils that inform the paleoanthropological record are primarily found preserved in caves and sinkholes. The purpose of this study is to assess the utility of Random Forest (RF) models for cave and sinkhole prediction in the Cradle of Humankind, South Africa. Multispectral satellite imagery, digital elevation models (DEMs), and geologic maps were converted into raster (pixelated matrix) images in a GIS environment to denote varying aspects of the local topography, including elevation, slope, aspect, curvature, drainage, spectral reflectance, vegetation cover, fault proximity, and underlying geology. The rasters were stacked and overlaid with 1080 known cave and sinkhole locality points and 1080 random non-cave points in the study area for model training. Variable values associated with these geopoints were input into an RF model in Python for training and evaluation using a spatial ten-fold cross-validation. The model performed with 81.6% accuracy and an area under the curve (AUC) of 0.912. The importance of each variable for prediction was evaluated by measuring the increase in prediction error when variable values were shuffled. Distance to major faults, location within the Chuniespoort geologic group, dolomite presence, chert presence, and elevation exhibited the highest importance for model accuracy, while three out of 48 total predictor variables exhibited less importance than a randomly generated variable. The identification of important/unimportant variables will help build more efficient, robust models in future iterations, as well as help identify variables that could be useful in other karst regions.
- Peer Review Report
- 10.7554/elife.78491.sa0
- Sep 5, 2022
Clinical prediction rules could help identify children at risk of slowed growth after an episode of diarrheal illness, and these rules may be generalizable to all children, regardless of diarrhea status.
- Peer Review Report
- 10.7554/elife.78491.sa1
- Sep 5, 2022
Clinical prediction rules could help identify children at risk of slowed growth after an episode of diarrheal illness, and these rules may be generalizable to all children, regardless of diarrhea status.
- Research Article
149
- 10.3390/ijerph17124206
- Jun 1, 2020
- International Journal of Environmental Research and Public Health
To compare the random forest (RF) model and the frequency ratio (FR) model for landslide susceptibility mapping (LSM), this research selected Yunyang Country as the study area for its frequent natural disasters; especially landslides. A landslide inventory was built by historical records; satellite images; and extensive field surveys. Subsequently; a geospatial database was established based on 987 historical landslides in the study area. Then; all the landslides were randomly divided into two datasets: 70% of them were used as the training dataset and 30% as the test dataset. Furthermore; under five primary conditioning factors (i.e., topography factors; geological factors; environmental factors; human engineering activities; and triggering factors), 22 secondary conditioning factors were selected to form an evaluation factor library for analyzing the landslide susceptibility. On this basis; the RF model training and the FR model mathematical analysis were performed; and the established models were used for the landslide susceptibility simulation in the entire area of Yunyang County. Next; based on the analysis results; the susceptibility maps were divided into five classes: very low; low; medium; high; and very high. In addition; the importance of conditioning factors was ranked and the influence of landslides was explored by using the RF model. The area under the curve (AUC) value of receiver operating characteristic (ROC) curve; precision; accuracy; and recall ratio were used to analyze the predictive ability of the above two LSM models. The results indicated a difference in the performances between the two models. The RF model (AUC = 0.988) performed better than the FR model (AUC = 0.716). Moreover; compared with the FR model; the RF model showed a higher coincidence degree between the areas in the high and the very low susceptibility classes; on the one hand; and the geographical spatial distribution of historical landslides; on the other hand. Therefore; it was concluded that the RF model was more suitable for landslide susceptibility evaluation in Yunyang County; because of its significant model performance; reliability; and stability. The outcome also provided a theoretical basis for application of machine learning techniques (e.g., RF) in landslide prevention; mitigation; and urban planning; so as to deliver an adequate response to the increasing demand for effective and low-cost tools in landslide susceptibility assessments.
- Research Article
4
- 10.1016/j.rcsop.2023.100307
- Jul 10, 2023
- Exploratory Research in Clinical and Social Pharmacy
Assessing treatment switch among patients with multiple sclerosis: A machine learning approach
- Research Article
46
- 10.1016/j.jjcc.2021.06.002
- Jun 19, 2021
- Journal of Cardiology
Predicting 30-day mortality after ST elevation myocardial infarction: Machine learning- based random forest and its external validation using two independent nationwide datasets
- Abstract
- 10.1136/ijgc-2021-igcs.167
- Nov 1, 2021
- International Journal of Gynecologic Cancer
ObjectivesTo train various machine learning algorithms to predict recurrence and recurrence-free survival (RFS) in high-grade endometrial cancer (HGEC)MethodsData was retrospectively collected across 8 Canadian centers including 1237 patients and divided...
- Research Article
- 10.1016/j.slast.2026.100436
- Jul 1, 2026
- SLAS technology
This study was to optimize the current methods for identifying and predicting the risk of critical illness in patients with connective tissue disease-associated interstitial lung disease (CTD-ILD). First, 200 patients diagnosed with CTD-ILD were included, and detailed demographic, serological, and imaging data were collected. Second, a risk identification and prediction framework was constructed based on multivariate logistic regression and machine learning algorithms (random forest (RF) and convolutional neural network (CNN)) to identify significant determinants of critical illness. Finally, the overall performance of each model was evaluated using K-fold cross-validation and external validation procedures. A feature ablation experiment was conducted based on the optimal random forest model to validate the independent contribution of each core predictor. The results showed that the logistic regression, random forest (RF), and CNN models were all successfully constructed and validated, among which the RF model demonstrated the best overall performance, with an accuracy of 85.7%, an area under the curve (AUC) of 0.88, a sensitivity of 83.5%, and a specificity of 88.2%. The ablation experiment confirmed that each feature had independent predictive value, with the most significant decline in model performance observed after the removal of IL‑6. Among them, the individual AUC value of interleukin-6 (IL-6) reached 0.981. Significant risk factors included patient age, C-reactive protein (CRP) level, presence of honeycomb lung on imaging, and the ratio of arterial oxygen partial pressure to inhaled oxygen concentration (PaO2/FiO2). The model in this study demonstrated satisfactory predictive ability and stability in both internal and external validation phases. The random forest model performed excellently in predicting the likelihood of critical illness in patients with CTD-ILD.
- Research Article
6
- 10.21037/atm-22-3049
- Dec 1, 2022
- Annals of Translational Medicine
BackgroundRadiation pneumonitis (RP) is a type of toxicity commonly associated with thoracic radiation therapy. We sought to establish a random forest (RF) model and evaluate its ability to predict RP in patients with non-small cell lung cancer (NSCLC) receiving moderately hypofractionated radiotherapy (hypo-RT).MethodsA total of 106 patients with stage II–IVa NSCLC who received moderately hypofractionated helical tomotherapy (2.3–3.0 Gy/fraction) at Zhongshan Hospital were included. All enrolled patients were divided chronologically into the training (67 patients) and validation (39 patients) groups. Higher than or equal to grade 2 RP was defined as the end point. Logistic regression and RF models were established and compared using the receiver operating characteristic (ROC) and a confusion matrix in the training and validation groups.ResultsThe cumulative incidence of the end point was 25.4% and 17.9% in the training and validation groups, respectively. Logistic regression models were constructed by dosage parameters of total lungs, ipsilateral or contralateral lungs, respectively. ROC analysis revealed that the dosimetric factors of total lungs yielded a superior classification performance than did that of the ipsilateral or contralateral lungs [area under the curve (AUC) =0.920, AUC =0.701, and AUC =0.661, respectively]. Furthermore, the RF model yielded a better prediction capacity than did the traditional logistic model based on the dosimetric factors of the total lungs (accuracy: 88.06%; precision: 84.62%; sensitivity: 64.71%; specificity: 96.00%). Moreover, the RF identified mean lung dose [MLD; mean decrease gini (MDG) =5.74], V20 (MDG =4.62), and V35 (MDG =3.08) of total lungs as the most common primary differentiators of RP.ConclusionsOur RF model established based on the dosimetric parameters of the total lungs could accurately predict the RP risk in patients with NSCLC treated with moderately hypofractionated tomotherapy.
- Research Article
21
- 10.12998/wjcc.v9.i29.8729
- Oct 16, 2021
- World Journal of Clinical Cases
BACKGROUNDHypotension after the induction of anesthesia is known to be associated with various adverse events. The involvement of a series of factors makes the prediction of hypotension during anesthesia quite challenging.AIMTo explore the ability and effectiveness of a random forest (RF) model in the prediction of post-induction hypotension (PIH) in patients undergoing cardiac surgery.METHODSPatient information was obtained from the electronic health records of the Second Affiliated Hospital of Hainan Medical University. The study included patients, ≥ 18 years of age, who underwent cardiac surgery from December 2007 to January 2018. An RF algorithm, which is a supervised machine learning technique, was employed to predict PIH. Model performance was assessed by the area under the curve (AUC) of the receiver operating characteristic. Mean decrease in the Gini index was used to rank various features based on their importance.RESULTSOf the 3030 patients included in the study, 1578 (52.1%) experienced hypotension after the induction of anesthesia. The RF model performed effectively, with an AUC of 0.843 (0.808-0.877) and identified mean blood pressure as the most important predictor of PIH after anesthesia. Age and body mass index also had a significant impact.CONCLUSIONThe generated RF model had high discrimination ability for the identification of individuals at high risk for a hypotensive event during cardiac surgery. The study results highlighted that machine learning tools confer unique advantages for the prediction of adverse post-anesthesia events.
- Research Article
4
- 10.3389/fneur.2025.1550789
- Apr 7, 2025
- Frontiers in Neurology
Background and aimParkinson’s disease (PD) is a neurodegenerative disorder with significant variability in disease progression. Identifying clinical and environmental risk factors associated with severe progression is essential for early diagnosis and personalized treatment. This study evaluates the performance of Random Forest (RF) and Logistic Regression (LR) models in forecasting the major risk factors associated with severe PD progression.MethodsWe performed a retrospective analysis of 378 PD patients (aged 40–75 years) with at 2 years of follow-up. The dataset included patient demographics, clinical features, medication history, comorbidities, and environmental exposures. The data were randomly split into a training group (70%) and a validation group (30%). Both the RF and LR models were trained on the training set, and performance was assessed through accuracy, sensitivity, specificity, and the Area Under the Curve (AUC) derived from ROC analysis.ResultsBoth models identified similar risk factors for severe PD progression, including older age, tremor-dominant motor subtype, long-term levodopa use, comorbid depression, and occupational pesticide exposure. The RF model outperformed the LR model, achieving an AUC of 0.85, accuracy of 82%, sensitivity of 79%, and specificity of 85%. In comparison, the LR model had an AUC of 0.78, accuracy of 76%, sensitivity of 74%, and specificity of 79%. ROC analysis showed that while both models could distinguish between slow and rapid disease progression, the RF model had stronger discriminatory power, particularly for identifying high-risk patients.ConclusionThe RF model provides better predictive accuracy and discriminatory power compared to Logistic Regression in identifying risk factors for severe PD progression. This study highlights the potential of machine learning techniques like Random Forest for early risk stratification and personalized management of PD.
- Research Article
- 10.1186/s12889-026-26420-6
- Feb 6, 2026
- BMC Public Health
From a public health perspective, the relationship between individual iodine nutritional and its associated risk factors has not been fully elucidated. The aim of this study is to utilize multiple biomarkers to represent individual iodine nutritional status, identify contributing factors for iodine imbalance, and develop a predictive assessment model for iodine nutrition evaluation in different water iodine districts. A total of 2,692 participants were recruited from Shandong and Anhui provinces in China. The study population was initially stratified into high-iodine and low-iodine groups based on water iodine concentrations of their residence. Thyroid function indicators and thyroid volume were used as assessment parameters. Both studies first utilized univariate regression to screen variables. After filtering out noisy features, the remaining significant variables were used to split the data into training and testing sets at a 7:3 ratio. Using random forest and eXtreme Gradient Boosting (XGBoost) models, we analyzed how modifiable factors (diet, medical history, lifestyle) relate to iodine homeostasis. Model performance was validated on the testing sets, with accuracy, sensitivity, and area under the curve (AUC) as key metrics. Integrated analysis of univariate regression, random forest, and XGBoost models revealed significant associations between drinking water sources and disrupted iodine homeostasis. In high-iodine areas, the XGBoost model demonstrated exceptional predictive performance for thyroid volume (R²=0.98, RMSE = 3.53). The results of the random forest classification model showed that the AUC was 0.76 (95% CI: 0.68–0.85) when TSH was used as the assessment indicator, while the AUC for TGAb and TPOAb were 0.74 (95% CI: 0.63–0.84) and 0.67 (95% CI: 0.54–0.80), respectively. In iodine-deficient areas, the XGBoost model maintained good predictive ability for thyroid volume (R2 = 0.97, RMSE = 3.28). The random forest model demonstrated moderate diagnostic accuracy among the biomarkers: TSH (AUC = 0.66; 95% CI: 0.51–0.80), TGAb (AUC = 0.69; 95% CI: 0.55–0.83), and TPOAb (AUC = 0.69; 95% CI: 0.56–0.83). This study established an individualized iodine nutrition assessment model by integrating multi-dimensional biochemical indicators with advanced machine learning algorithms. The model represented by thyroid volume effectively identified the key factors disrupting iodine homeostasis and was capable of accurately predicting individual iodine nutritional status. Its dual utility provides: (1) evidence-based quantitative metrics that can offer personalized guidance for iodine supplementation in clinical practice; and (2) a decision-support framework for regions with varying iodine levels, which can inform the optimization of iodine supplementation programs in specific areas.
- Research Article
100
- 10.3389/feart.2021.589630
- Mar 31, 2021
- Frontiers in Earth Science
Landslide susceptibility mapping is very important for landslide risk evaluation and land use planning. Toward this end, this paper presents a case study in Ningqiang County, Shanxi Province, China. Slope units were selected as the basic mapping units. A traditional statistical certainty factor model (CF), a machine learning support vector machine model (SVM) and random forest model (RF), along with a hybrid CF-SVM model and a CF-RF model were applied to analyze landslide susceptibility. Firstly, 10 landslide conditioning factors were selected, namely slope-angle, altitude, slope aspect, degree of relief, lithology, distance to rivers, distance to faults, distance to roads, average annual rainfall and normalized difference vegetation index. The 23,169 slope units were generated from a Digital Elevation Model and the corresponding 10 conditioning factor layers were produced from both geological and geographical data. Then, landslide susceptibility mapping was carried out using the five models, respectively. Next, the landslide density (LD), frequency ratio (FR), the area under the curve (AUC) and other indicators were used to validate the rationality, performance and accuracy of the models. The results showed that the susceptibility maps produced from the different models were all reasonable. In each map, the LD and FR were greatest in the zones classed as having very high landslide susceptibility, followed by the high, moderate, low and very low landslide susceptibility classes, respectively. From the comparison of the different maps and ROC curves, the RF model based on slope units was the most appropriate for landslide susceptibility mapping in the study area. It was also found that the combination of weaker learner model (CF model here) with a stronger learner model (SVM and RF model here) can impact the applicability of the stronger model.
- Research Article
115
- 10.1016/j.catena.2018.12.013
- Dec 13, 2018
- CATENA
Susceptibility assessment of landslides triggered by earthquakes in the Western Sichuan Plateau
- Research Article
36
- 10.1016/j.nhres.2023.07.004
- Jul 20, 2023
- Natural Hazards Research
Comparative study on landslide susceptibility mapping based on different ratios of training samples and testing samples by using RF and FR-RF models
- Research Article
34
- 10.1080/01431161.2018.1430399
- Jan 25, 2018
- International Journal of Remote Sensing
ABSTRACTMapping the spatial distribution of soil classes is important for informing soil use and management decisions. This study aimed to effectively implement Random Forest (RF) model and to evaluate the behaviour and performance of the model for soil classification of Indian districts. Soil-forming factors, known as ‘scorpan,’ are selected as environmental covariates to tune RF model to classify 11 different soil categories. Thirty-five digital layers are prepared using different satellite data [ALOS (Advanced Land Observing Satellite) digital elevation model, Landsat-8, Moderate Resolution Imaging Spectroradiometer normalized difference vegetation index product, RISAT-1 (Radar Imaging Satellite-1), Sentinel-1A] and climatic data (precipitation and temperature) to represent scorpan environmental covariates in the study area. The RF parameters corresponding to highest Cohen’s kappa coefficient (κ) value and lowest number of random split variables are considered optimum values for RF model. Model behaviour evaluation is based on mapping accuracy, sensitivity to data set size, and noise. Two other machine-learning methods, CART (Classification and Regression Tree) decision tree (CDT) and CART ensemble bagger (CEB), are used to provide the comparative study. To access behaviour of models to the false data set, noise in training set is produced by assigning a false class to the training set in 5% increment. Comparative performance of RF model is based on quality assessment measures. To evaluate the performance of models, marginal rates, F-measure, and Jaccard’s coefficient of the community, classification success index and agreement coefficients are selected under quality assessment measures. The score is calculated to rank the algorithm. RF model shows high stability against data set reduction in comparison to other methods. The results show that the abrupt change in accuracy is only observed after 60% training data reduction in RF model; however, significant decrease in accuracy can be noted after 45% and 25% data reduction in CEB and CDT, respectively. The RF model shows comparatively the greater resistance to noise. Overall, RF model has performed better than CDT and CEB to classify soil categories in the study area. The results of this research provide new insights into the performance of RF in the context of soil class mapping.