A novel methodological framework for predicting and mapping agriculture-related soil attributes using Euclidean distance, regular grids, and machine learning algorithms
Recent advances in statistical and machine learning (ML) methods have improved the prediction of soil attributes at fine spatial scales, yet the comparative performance and reliability of these techniques remain unclear. This study compared Ordinary Kriging (OK), Inverse Distance Weighting (IDW), and ML algorithms in predicting and spatializing soil attributes, while also evaluating prediction uncertainty and computational processing time. Conducted in Minas Gerais State (Brazil), the analysis used Euclidean distance based predictors derived from X-Y coordinates and regular grids with 5, 7, and 10 divisions. Soil attribute maps (CEC, phosphorus, sand, and clay) were generated using OK, IDW, Random Forest (RF), Cubist, Support Vector Machine (SVM), and Earth. Model performance was assessed using R2, RMSE, MAE, and the coefficient of variation. IDW and OK showed the lowest predictive accuracy (R2 = 0.52–0.58), whereas ML methods, especially RF and SVM achieved superior performance (R2 = 0.62–0.70). Among ML algorithms, Earth performed worst, while RF produced the highest accuracy for all attributes except sand, for which SVM performed best. Processing time was shortest for IDW, followed by OK; among ML models, Earth was fastest, followed by RF, SVM, and Cubist. Larger regular grids improved ML prediction and spatialization but increased computational cost. ML methods thus outperform traditional geostatistical interpolators, benefiting from the use of numerous covariates and flexible algorithmic structures, although requiring greater computational time. These findings demonstrate the robustness and practical potential of ML approaches for soil attribute mapping.
- Research Article
- 10.1002/fsat.3304_6.x
- Dec 1, 2019
- Food Science and Technology
Sensors support machine learning
- Research Article
14
- 10.1186/s41043-024-00647-8
- Oct 12, 2024
- Journal of Health, Population and Nutrition
Background and aimsThe birth weight of a newborn is a crucial factor that affects their overall health and future well-being. Low birth weight (LBW) is a widespread global issue, which the World Health Organization defines as weighing less than 2,500 g. LBW can have severe negative consequences on an individual’s health, including neonatal mortality and various health concerns throughout their life. To address this problem, this study has been conducted using BDHS 2017–2018 data to uncover important aspects of LBW using a variety of machine learning (ML) approaches and to determine the best feature selection technique and best predictive ML model.MethodsTo pick out the key features, the Boruta algorithm and wrapper method were used. Logistic Regression (LR) used as traditional method and several machine learning classifiers were then used, including, DT (Decision Tree), SVM (Support Vector Machine), NB (Naïve Bayes), RF (Random Forest), XGBoost (eXtreme Gradient Boosting), and AdaBoost (Adaptive Boosting), to determine the best model for predicting LBW. The model’s performance was evaluated based on the specificity, sensitivity, accuracy, F1 score and AUC value.ResultsResult shows, Boruta algorithm identifies eleven significant features including respondent’s age, highest education level, educational attainment, wealth index, age at first birth, weight, height, BMI, age at first sexual intercourse, birth order number, and whether the child is a twin. Incorporating Boruta algorithm’s significant features, the performance of traditional LR and ML methods including DT, SVM, NB, RF, XGBoost, and AB were evaluated where LR, had a specificity, sensitivity, accuracy and F1 score of 0.85, 0.5, 85.15% and 0.915. While the ML methods DT, SVM, NB, RF, XGBoost, and AB model’s respective accuracy values were 85.35%, 85.15%, 84.54%, 81.18%, and 84.41%. Based on the specificity, sensitivity, accuracy, F1 score and AUC, RF (specificity = 0.99, sensitivity = 0.58, accuracy = 85.86%, F1 score = 0.9243, AUC = 0.549) outperformed the other methods. Both the classical (LR) and machine learning (ML) models’ performance has improved dramatically when important characteristics are extracted using the wrapper method. The LR method identified five significant features with a specificity, sensitivity, accuracy and F1 score of 0.87, 0.33, 87.12% and 0.9309. The region, whether the infant is a twin, and cesarean delivery were the three key features discovered by the DT and RF models, which were implemented using the wrapper technique. All three models had the identical F1 score of 0.9318. However, “child is twin” was recognized as a significant feature by the SVM, NB, and AB models, with an F1 score of 0.9315. Ultimately, with an F1 score of 0.9315, the XGBoost model recognized “child is twin” and “age at first sex” as relevant features. Random Forest again beat the other approaches in this instance.ConclusionsThe study reveals Wrapper method as the optimal feature selection technique. The ML method outperforms traditional methods, with Random Forest (RF) being the most effective predictive model for Low-Birth-Weight prediction. The study suggests that policymakers in Bangladesh can mitigate low birth weight newborns by considering identified risk factors.
- Abstract
- 10.1016/j.spinee.2021.05.333
- Aug 10, 2021
- The Spine Journal
P125. Development of a novel ensemble machine learning algorithm for prediction of complications and readmission after anterior cervical spinal fusion
- Abstract
- 10.1016/j.spinee.2021.05.334
- Aug 10, 2021
- The Spine Journal
P126. Development of a novel ensemble machine learning algorithm for prediction of complications and readmission after posterior cervical spinal fusion
- Abstract
- 10.1016/j.jval.2020.04.1006
- May 1, 2020
- Value in Health
PND117 IDENTIFYING PREDICTORS OF HIGH-COST MULTIPLE SCLEROSIS PATIENTS: A MACHINE LEARNING APPROACH
- Research Article
750
- 10.1139/er-2020-0019
- Jul 28, 2020
- Environmental Reviews
Artificial intelligence has been applied in wildfire science and management since the 1990s, with early applications including neural networks and expert systems. Since then, the field has rapidly progressed congruently with the wide adoption of machine learning (ML) methods in the environmental sciences. Here, we present a scoping review of ML applications in wildfire science and management. Our overall objective is to improve awareness of ML methods among wildfire researchers and managers, as well as illustrate the diverse and challenging range of problems in wildfire science available to ML data scientists. To that end, we first present an overview of popular ML approaches used in wildfire science to date and then review the use of ML in wildfire science as broadly categorized into six problem domains, including (i) fuels characterization, fire detection, and mapping; (ii) fire weather and climate change; (iii) fire occurrence, susceptibility, and risk; (iv) fire behavior prediction; (v) fire effects; and (vi) fire management. Furthermore, we discuss the advantages and limitations of various ML approaches relating to data size, computational requirements, generalizability, and interpretability, as well as identify opportunities for future advances in the science and management of wildfires within a data science context. In total, to the end of 2019, we identified 300 relevant publications in which the most frequently used ML methods across problem domains included random forests, MaxEnt, artificial neural networks, decision trees, support vector machines, and genetic algorithms. As such, there exists opportunities to apply more current ML methods — including deep learning and agent-based learning — in the wildfire sciences, especially in instances involving very large multivariate datasets. We must recognize, however, that despite the ability of ML models to learn on their own, expertise in wildfire science is necessary to ensure realistic modelling of fire processes across multiple scales, while the complexity of some ML methods such as deep learning requires a dedicated and sophisticated knowledge of their application. Finally, we stress that the wildfire research and management communities play an active role in providing relevant, high-quality, and freely available wildfire data for use by practitioners of ML methods.
- Research Article
96
- 10.3390/s20185055
- Sep 5, 2020
- Sensors (Basel, Switzerland)
The vegetation index (VI) has been successfully used to monitor the growth and to predict the yield of agricultural crops. In this paper, a long-term observation was conducted for the yield prediction of maize using an unmanned aerial vehicle (UAV) and estimations of chlorophyll contents using SPAD-502. A new vegetation index termed as modified red blue VI (MRBVI) was developed to monitor the growth and to predict the yields of maize by establishing relationships between MRBVI- and SPAD-502-based chlorophyll contents. The coefficients of determination (R2s) were 0.462 and 0.570 in chlorophyll contents’ estimations and yield predictions using MRBVI, and the results were relatively better than the results from the seven other commonly used VI approaches. All VIs during the different growth stages of maize were calculated and compared with the measured values of chlorophyll contents directly, and the relative error (RE) of MRBVI is the lowest at 0.355. Further, machine learning (ML) methods such as the backpropagation neural network model (BP), support vector machine (SVM), random forest (RF), and extreme learning machine (ELM) were adopted for predicting the yields of maize. All VIs calculated for each image captured during important phenological stages of maize were set as independent variables and the corresponding yields of each plot were defined as dependent variables. The ML models used the leave one out method (LOO), where the root mean square errors (RMSEs) were 2.157, 1.099, 1.146, and 1.698 (g/hundred grain weight) for BP, SVM, RF, and ELM. The mean absolute errors (MAEs) were 1.739, 0.886, 0.925, and 1.356 (g/hundred grain weight) for BP, SVM, RF, and ELM, respectively. Thus, the SVM method performed better in predicting the yields of maize than the other ML methods. Therefore, it is strongly suggested that the MRBVI calculated from images acquired at different growth stages integrated with advanced ML methods should be used for agricultural- and ecological-related chlorophyll estimation and yield predictions.
- Research Article
16
- 10.1016/j.envpol.2022.120931
- Dec 21, 2022
- Environmental Pollution
Three-dimensional spatial prediction of Zn in the soil of a former tire manufacturing plant using machine learning and readily attainable multisource auxiliary data
- Research Article
56
- 10.3390/su15065341
- Mar 17, 2023
- Sustainability
Air pollution in Macau has become a serious problem following the Pearl River Delta’s (PRD) rapid industrialization that began in the 1990s. With this in mind, Macau needs an air quality forecast system that accurately predicts pollutant concentration during the occurrence of pollution episodes to warn the public ahead of time. Five different state-of-the-art machine learning (ML) algorithms were applied to create predictive models to forecast PM2.5, PM10, and CO concentrations for the next 24 and 48 h, which included artificial neural networks (ANN), random forest (RF), extreme gradient boosting (XGBoost), support vector machine (SVM), and multiple linear regression (MLR), to determine the best ML algorithms for the respective pollutants and time scale. The diurnal measurements of air quality data in Macau from 2016 to 2021 were obtained for this work. The 2020 and 2021 datasets were used for model testing, while the four-year data before 2020 and 2021 were used to build and train the ML models. Results show that the ANN, RF, XGBoost, SVM, and MLR models were able to provide good performance in building up a 24-h forecast with a higher coefficient of determination (R2) and lower root mean square error (RMSE), mean absolute error (MAE), and biases (BIAS). Meanwhile, all the ML models in the 48-h forecasting performance were satisfactory enough to be accepted as a two-day continuous forecast even if the R2 value was lower than the 24-h forecast. The 48-h forecasting model could be further improved by proper feature selection based on the 24-h dataset, using the Shapley Additive Explanations (SHAP) value test and the adjusted R2 value of the 48-h forecasting model. In conclusion, the above five ML algorithms were able to successfully forecast the 24 and 48 h of pollutant concentration in Macau, with the RF and SVM models performing the best in the prediction of PM2.5 and PM10, and CO in both 24 and 48-h forecasts.
- Research Article
65
- 10.1186/s12903-021-01996-0
- Dec 1, 2021
- BMC Oral Health
BackgroundRecently, the dental age estimation method developed by Cameriere has been widely recognized and accepted. Although machine learning (ML) methods can improve the accuracy of dental age estimation, no machine learning research exists on the use of the Cameriere dental age estimation method, making this research innovative and meaningful.AimThe purpose of this research is to use 7 lower left permanent teeth and three models [random forest (RF), support vector machine (SVM), and linear regression (LR)] based on the Cameriere method to predict children's dental age, and compare with the Cameriere age estimation.Subjects and methodsThis was a retrospective study that collected and analyzed orthopantomograms of 748 children (356 females and 392 males) aged 5–13 years. Data were randomly divided into training and test datasets in an 80–20% proportion for the ML algorithms. The procedure, starting with randomly creating new training and test datasets, was repeated 20 times. 7 permanent developing teeth on the left mandible (except wisdom teeth) were recorded using the Cameriere method. Then, the traditional Cameriere formula and three models (RF, SVM, and LR) were used to estimate the dental age. The age prediction accuracy was measured by five indicators: the coefficient of determination (R2), mean error (ME), root mean square error (RMSE), mean square error (MSE), and mean absolute error (MAE).ResultsThe research showed that the ML models have better accuracy than the traditional Cameriere formula. The ME, MAE, MSE, and RMSE values of the SVM model (0.004, 0.489, 0.392, and 0.625, respectively) and the RF model (− 0.004, 0.495, 0.389, and 0.623, respectively) were lower with the highest accuracy. In contrast, the ME, MAE, MSE and RMSE of the European Cameriere formula were 0.592, 0.846, 0.755, and 0.869, respectively, and those of the Chinese Cameriere formula were 0.748, 0.812, 0.890 and 0.943, respectively.ConclusionsCompared to the Cameriere formula, ML methods based on the Cameriere’s maturation stages were more accurate in estimating dental age. These results support the use of ML algorithms instead of the traditional Cameriere formula.
- Research Article
- 10.3390/a18080482
- Aug 4, 2025
- Algorithms
Simulated data created in silico using a previously reported method were sampled by bootstrapping to generate data sets for training multiple copies of an ensemble learner (i.e., a machine learning (ML) method). The posterior probabilities of class membership obtained by applying the ensemble of ML models to previously unseen validation data were fitted to a beta distribution. The shape parameters for the fitted distribution were used to calculate the subjective opinion of sample membership into one of two mutually exclusive classes. The subjective opinion consists of belief, disbelief and uncertainty masses. A subjective opinion for each validation sample allows identification of high-uncertainty predictions. The projected probabilities of the validation opinions were used to calculate log-likelihood ratio scores and generate receiver operating characteristic (ROC) curves from which an opinion-supported decision can be made. Three very different ML models, linear discriminant analysis (LDA), random forest (RF), and support vector machines (SVM) were applied to the two-state classification problem in the analysis of forensic fire debris samples. For each ML method, a set of 100 ML models was trained on data sets bootstrapped from 60,000 in silico samples. The impact of training data set size on opinion uncertainty and ROC area under the curve (AUC) were studied. The median uncertainty for the validation data was smallest for LDA ML and largest for the SVM ML. The median uncertainty continually decreased as the size of the training data set increased for all ML.The AUC for ROC curves based on projected probabilities was largest for the RF model and smallest for the LDA method. The ROC AUC was statistically unchanged for LDA at training data sets exceeding 200 samples; however, the AUC increased with increasing sample size for the RF and SVM methods. The SVM method, the slowest to train, was limited to a maximum of 20,000 training samples. All three ML methods showed increasing performance when the validation data was limited to higher ignitable liquid contributions. An ensemble of 100 RF ML models, each trained on 60,000 in silico samples, performed the best with a median uncertainty of 1.39x10−2 and ROC AUC of 0.849 for all validation samples.
- Research Article
33
- 10.1016/j.oregeorev.2024.105887
- Jan 13, 2024
- Ore Geology Reviews
Deposit type discrimination based on trace elements in sphalerite
- Research Article
40
- 10.3390/ijgi9040276
- Apr 23, 2020
- ISPRS International Journal of Geo-Information
In the current paper we assess different machine learning (ML) models and hybrid geostatistical methods in the prediction of soil pH using digital elevation model derivates (environmental covariates) and co-located soil parameters (soil covariates). The study was located in the area of Grevena, Greece, where 266 disturbed soil samples were collected from randomly selected locations and analyzed in the laboratory of the Soil and Water Resources Institute. The different models that were assessed were random forests (RF), random forests kriging (RFK), gradient boosting (GB), gradient boosting kriging (GBK), neural networks (NN), and neural networks kriging (NNK) and finally, multiple linear regression (MLR), ordinary kriging (OK), and regression kriging (RK) that although they are not ML models, they were used for comparison reasons. Both the GB and RF models presented the best results in the study, with NN a close second. The introduction of OK to the ML models’ residuals did not have a major impact. Classical geostatistical or hybrid geostatistical methods without ML (OK, MLR, and RK) exhibited worse prediction accuracy compared to the models that included ML. Furthermore, different implementations (methods and packages) of the same ML models were also assessed. Regarding RF and GB, the different implementations that were applied (ranger-ranger, randomForest-rf, xgboost-xgbTree, xgboost-xgbDART) led to similar results, whereas in NN, the differences between the implementations used (nnet-nnet and nnet-avNNet) were more distinct. Finally, ML models tuned through a random search optimization method were compared with the same ML models with their default values. The results showed that the predictions were improved by the optimization process only where the ML algorithms demanded a large number of hyperparameters that needed tuning and there was a significant difference between the default values and the optimized ones, like in the case of GB and NN, but not in RF. In general, the current study concluded that although RF and GB presented approximately the same prediction accuracy, RF had more consistent results, regardless of different packages, different hyperparameter selection methods, or even the inclusion of OK in the ML models’ residuals.
- Abstract
3
- 10.1182/blood-2023-182438
- Nov 2, 2023
- Blood
Machine Learning Validates Risk Biomarkers of Chronic Graft-Versus-Host Disease in 936 Patients from BMT CTN 0201 & 1202 Cohorts
- Research Article
3
- 10.2174/0118741495343680240911053413
- Oct 7, 2024
- The Open Civil Engineering Journal
Aim This study aims to enhance safety in large diameter tunnel construction by integrating robust optimization and machine learning (ML) techniques with Building Information Modeling (BIM). By acquiring and preprocessing various datasets, implementing feature engineering, and using algorithms like SVM, decision trees, ANN, and random forests, the study demonstrates the effectiveness of ML models in risk prediction and mitigation, ultimately advancing safety performance in civil engineering projects. Background Large diameter tunnel construction presents significant safety challenges. Traditional methods often fall short of effectively predicting and mitigating risks. This study addresses these gaps by integrating robust optimization and machine learning (ML) approaches with Building Information Modeling (BIM) technology. By acquiring and preprocessing diverse datasets, implementing feature engineering, and employing ML algorithms, the study aims to enhance risk prediction and safety measures in tunnel construction projects. Objective The objective of this study is to improve safety in large diameter tunnel construction by integrating robust optimization and machine learning (ML) techniques with Building Information Modeling (BIM). This involves acquiring and preprocessing diverse datasets, using feature engineering to extract key parameters, and applying ML algorithms like SVM, decision trees, ANN, and random forests to predict and mitigate risks, ultimately enhancing safety performance in civil engineering projects. Methods The study's methods include acquiring and preprocessing various datasets (geological, structural, environmental, operational, historical, and simulation). Feature engineering techniques are used to extract key safety parameters for tunnels. Machine learning algorithms, such as decision trees, support vector machines (SVM), artificial neural networks, and random forests, are employed to analyze the data and predict construction risks. The SVM algorithm, with a 98.76% accuracy, is the most reliable predictor. Results The study found that the Support Vector Machine (SVM) algorithm was the most accurate predictor of risks in large diameter tunnel construction, achieving a 98.76% accuracy rate. Other models, such as decision trees, artificial neural networks, and random forests, also performed well, validating the effectiveness of ML-based solutions for risk assessment and mitigation. These predictive models enable stakeholders to monitor construction, allocate resources, and implement preventative measures effectively. Conclusion The study concludes that integrating machine learning (ML) approaches with Building Information Modeling (BIM) significantly improves safety in large diameter tunnel construction. The Support Vector Machine (SVM) algorithm, with 98.76% accuracy, is the most reliable predictor of risks. Other models, like decision trees, artificial neural networks, and random forests, also perform well, validating ML-based solutions for risk assessment. Adopting these ML approaches enhances safety performance and resource management in civil engineering projects.