Nondestructive prediction of leaf area in colored cotton cultivars: a comparative approach using machine learning models with an interactive web interface
Abstract Background Leaf area is a crucial indicator of plant growth and physiology, with direct measurements being destructive to the plant. This study aimed to develop and compare machine learning models [support vector regression (SVR), adaptive neuro-fuzzy inference system (ANFIS), and deep multilayer perceptron (DMLP)] and linear regression (LRM) for the nondestructive prediction of leaf area in five colored cotton cultivars. A total of 1 334 leaves were sampled, and their length (L), width (W), and leaf area (LA) were determined via digitized images. The models were developed using 70% of the data for training and 30% for validation. Their performance was evaluated using the coefficient of determination ( R 2 ), root mean square error, mean absolute error, mean absolute percentage error, and Willmott's index of agreement. Results The results showed that the machine learning models, notably the ANFIS (triangular membership function), the DMLP (2–16-16–1 configuration), and the SVR [radial basis function (RBF) kernel], significantly outperformed the linear regression models in leaf area estimation accuracy. The ANFIS and DMLP models achieved the highest R 2 (0.979 3, test), followed by the SVR model ( R 2 = 0.979 0, test), all with minimal errors. Among the linear models, the LRM (using the L × W product) was the most effective ( R 2 = 0.978 3). Conclusions On the basis of the performance criteria of the models, the machine learning models are more accurate for the nondestructive estimation of leaf area in colored cotton. The best-performing model (SVR with RBF kernel) was made available in an interactive web application, aiming to optimize crop management with accurate and nondestructive data.
- Research Article
12
- 10.1155/2021/1519019
- Nov 11, 2021
- Journal of Mathematics
In recent years, as global financial markets have become increasingly connected, the degree of correlation between financial assets has become closer, and technological advances have made the transmission of information faster and faster, and information networks have integrated capital markets into one, making it easier for single financial market risk problems to form systemic risk through a high degree of market linkage effects. Based on the characteristics of financial markets containing both linear and nonlinear components, this paper chooses to use Autoregressive Integrated Moving Average (ARIMA) model and feedback Support Vector Regression (SVR) models to effectively integrate the ARIMA model and the SVR model, taking into account their respective linear and nonlinear characteristics. The paper chooses to use the (Autoregressive Integrated Moving Average (ARIMA) model and feedback Support Vector Regression (SVR) models to effectively integrate the strengths of the ARIMA and SVR models in terms of linearity and nonlinearity to perform forecasting analysis of financial markets. One of the important functions of forecasting is to transform future uncertainty into measurable risk, so that we can base our plans and actions on it. In this paper, the combined ARIMA-SVR model is compared with the single ARIMA model and SVR model in terms of the mean absolute error (MAE), root mean square error (RMSE), and mean absolute percentage error (MAPE), where MAE and RMSE measure the absolute error between the predicted and true values, and MAPE measures the relative error between the predicted and true values. and the relative error between the true value. The results show that the combined ARIMA-SVR model has a better forecasting effect and higher forecasting accuracy than the single ARIMA model and SVR model, and the SVR model has higher forecasting accuracy than the ARIMA model in forecasting financial markets.
- Research Article
19
- 10.1016/j.aej.2021.04.096
- Jun 6, 2021
- Alexandria Engineering Journal
Missing data imputation of MAGDAS-9’s ground electromagnetism with supervised machine learning and conventional statistical analysis models
- Research Article
1
- 10.36647/ciml/05.02.a002
- Jan 1, 2025
- Computational Intelligence and Machine Learning
The devastating earthquake that caused significant destruction in 11 provinces on February 6, 2023, has accelerated the growth rate of the cement sector. This rapid growth, coupled with increasing stock market activity, marks a golden age for the sector while emphasizing the critical need for accurate future production forecasting. Leveraging the predictive capabilities of machine learning algorithms, experiments were conducted using five years of production data from a cement factory in the Southeastern Anatolia region. The Support Vector Regression (SVR) model, an application of the Support Vector Machine (SVM) algorithm, was tested with RBF, linear, sigmoid, and polynomial kernels. Among these, the SVR model with the RBF kernel yielded the best performance across four evaluation metrics: Mean Squared Error (MSE): 0.002926, Root Mean Squared Error (RMSE): 0.054094, Mean Absolute Error (MAE): 0.048611, and Mean Absolute Percentage Error (MAPE): 0.052697. This paper highlights the effectiveness of SVR-RBF in providing reliable production forecasts for the cement industry and supporting strategic planning to address dynamic market demands.
- Research Article
21
- 10.1016/j.istruc.2021.12.054
- Dec 27, 2021
- Structures
Axial strength prediction of steel tube confined concrete columns using a hybrid machine learning model
- Research Article
1
- 10.3390/foods13233858
- Nov 29, 2024
- Foods (Basel, Switzerland)
The article demonstrates the Brix content of melon fruits grafted with different varieties of rootstock using Support Vector Regression (SVR) and Multiple Linear Regression (MLR) model approaches. The analysis yielded primary fruit biochemical measurements on the following rootstocks, Sphinx, Albatros, and Dinero: nitrogen, phosphorus, potassium, calcium, and magnesium. Established models were evaluated with Mean Absolute Error (MAE), Mean Absolute Percentage Error (MAPE), Mean Square Error (MSE), Root Mean Square Error (RMSE), and Coefficient of Determination (R2) metrics. In the test section, the results of the MLR model were calculated as MAE: 0.0728, MAPE: 0.0117, MSE: 0.0088, RMSE: 0.0936, and R2: 0.9472, while the results of the SVR model were calculated as MAE: 0.0334, MAPE: 0.0054, MSE: 0.0016, RMSE: 0.0398, and R2: 0.9904. Despite both models performing well, the SVR model showed superior accuracy, outperforming MLR by 54% to 82% in terms of predictions. The relationships between Brix levels and various nutrients, such as sucrose, glucose, and fructose, were found to be strong, while titratable acidity had a minimal effect. SVR was found to be a more reliable, non-destructive method for melon quality assessment. These findings revealed the relationship between Brix and sugar levels on melon quality. The study highlights the potential of these machine learning models in optimizing the rootstock effect and managing melon cultivation to improve fruit quality.
- Research Article
45
- 10.1016/j.scienta.2017.03.028
- Mar 28, 2017
- Scientia Horticulturae
Non-destructive estimation of leaf area of durian (Durio zibethinus) – An artificial neural network approach
- Research Article
35
- 10.3390/math10162971
- Aug 17, 2022
- Mathematics
Precise streamflow estimation plays a key role in optimal water resource use, reservoirs operations, and designing and planning future hydropower projects. Machine learning models were successfully utilized to estimate streamflow in recent years In this study, a new approach, covariance matrix adaptation evolution strategy (CMAES), was utilized to improve the accuracy of seven machine learning models, namely extreme learning machine (ELM), elastic net (EN), Gaussian processes regression (GPR), support vector regression (SVR), least square SVR (LSSVR), extreme gradient boosting (XGB), and radial basis function neural network (RBFNN), in predicting streamflow. The CMAES was used for proper tuning of control parameters of these selected machine learning models. Seven input combinations were decided to estimate streamflow based on previous lagged temperature and streamflow data values. For numerical prediction accuracy comparison of these machine learning models, six statistical indexes are used, i.e., relative root mean squared error (RRMSE), mean absolute error (MAE), mean absolute percentage error (MAPE), Nash–Sutcliffe efficiency (NSE), and the Kling–Gupta efficiency agreement index (KGE). In contrast, this study uses scatter plots, radar charts, and Taylor diagrams for graphically predicted accuracy comparison. Results show that SVR provided more accurate results than the other methods, especially for the temperature input cases. In contrast, in some streamflow input cases, the LSSVR and GPR were better than the SVR. The SVR tuned by CMAES with temperature and streamflow inputs produced the least RRMSE (0.266), MAE (263.44), and MAPE (12.44) in streamflow estimation. The EN method was found to be the worst model in streamflow prediction. Uncertainty analysis also endorsed the superiority of the SVR over other machine learning methods by having low uncertainty values. Overall, the SVR model based on either temperature or streamflow as inputs, tuned by CMAES, is highly recommended for streamflow estimation.
- Research Article
1
- 10.1088/1402-4896/ad9d03
- Dec 24, 2024
- Physica Scripta
This paper investigates the impact of weather and calendar data features on load profile analysis and proposes an improved load profile forecasting approach using Support Vector Regression (SVR). A detailed load profile was constructed for a single-family house in Algiers, Algeria, based on an in-depth analysis of survey responses over three years. The SVR model, employing a built-in split method, has demonstrated the highest efficiency for short-term predictions, particularly for one-day forecasts. The initial results from the standard SVR model yielded a Mean Absolute Percentage Error (MAPE) of 39.50%, a Mean Squared Error (MSE) of 0.0461, and an R 2 score of 0.8679. Additionally, the study compared the performance of other machine learning models, including Random Forest Regressor (RFR), Gradient Boosting Regressor (GBR), and Artificial Neural Network (ANN) for one-day forecasting. The RFR achieved an MSE of 0.20088, MAPE of 90.45%, and an R 2 score of 0.4243; the GBR yielded an MSE of 0.13274, MAPE of 80.76%, and an R 2 score of 0.6196; while the ANN demonstrated an MSE of 0.0618, MAPE of 59.71%, and an R 2 score of 0.6407. Notably, the SVR model emerged as the superior performer across various forecast horizons, prompting further exploration to enhance its capabilities. In addition to the standard SVR method, this study introduces an enhanced SVR approach utilizing the Radial Basis Function (RBF) kernel and fine-tuning its parameters. This enhanced model achieved a significantly reduced MSE of 0.0419, MAPE of 18.89%, and an improved R 2 score of 0.8799 for one-day forecasts, surpassing the standard SVR model's performance. Fourier transform analysis was also applied to uncover underlying patterns in the consumption data, complementing the time-domain results from the SVR model. A grid search optimized hyperparameters, revealing that C = 5 and ε = 0.01 provided the best model performance. These findings offer practical implications for energy management, policy-making, and the development of smart grid technologies, contributing to the sustainability and efficiency of energy consumption in residential settings.
- Research Article
1
- 10.3390/jcm14186373
- Sep 10, 2025
- Journal of Clinical Medicine
Background: Suicide remains a leading cause of death among youth, yet effective tools to predict suicide attempts (SA) in individuals under 18 are scarce. This study aims to develop machine learning (ML) models to predict SA in paediatric populations using Google Trends data. Methods: Relative Search Volumes (RSVs) from Google Trends were analysed for terms linked to suicide risk factors. Pearson Correlation Coefficients (PCC) identified terms strongly associated with SA rates. Based on these, several ML models were developed and evaluated, including Random Forest Regression, Support Vector Regression (SVR), XGBoost, and Linear Regression. Model performance was assessed using metrics such as PCC, mean absolute error (MAE), mean squared error (MSE), root mean square error (RMSE), and mean absolute percentage error (MAPE). Results: Terms related to suicide prevention and symptoms, including psychiatrist and anxiety disorder, showed the strongest correlations with SA rates (PCC ≥ 0.90). Random Forest Regression emerged as the top-performing ML model (PCC = 0.953, MAPE = 20.12%, RMSE = 17.21), highlighting burnout, anxiety disorder, antidepressants, and psychiatrist as key predictors of SA. Other models’ scores were XGBoost (PCC = 0.446, MAPE = 22.57%, RMSE = 18.03), SVR (PCC = 0.833, MAPE = 42.23%, RMSE = 47.32) and Linear Regression (PCC = 0.947, MAPE = 23.64%, RMSE = 17.66). Conclusions: Google Trends–based ML models suggest potential utility for short-term prediction of youth SA. These preliminary findings support the utility of search data in identifying real-time suicide risk in paediatric populations.
- Research Article
- 10.3390/met16030266
- Feb 27, 2026
- Metals
Research on eco-friendly and energy-efficient machining processes has gained significant importance within the domain of sustainable production. This study is focused on enhancing the energy performance and sustainability of the milling process. Four machine learning (ML) models, namely, multiple linear regression (MLR), support vector regression (SVR), Gaussian process regression (GPR), and adaptive network-based fuzzy inference system (ANFIS), were proposed to estimate specific energy consumption (SEC) in the milling of Ti6-Al4-V under two eco-benign cooling conditions: cryogenic and minimum quantity lubrication (MQL). Several statistical metrics, including normalized mean absolute error (nMAE), mean absolute percentage error (MAPE), normalized root mean square error (nRMSE), maximum absolute percentage error (maxAPE), coefficient of determination (R2), and Willmott’s index of agreement (IA), were employed to validate the performances of the ML models. A high level of agreement between the predicted and experimental SEC data for both the training and test datasets supports the reliability of the proposed ML models. Although the MLR model performed well, the results revealed that the other ML models demonstrated better overall performance. According to the statistical metrics, the models’ predictive performance improved in the following sequence: MLR, SVR, GPR, and finally ANFIS, which demonstrated the highest predictive capability.
- Research Article
13
- 10.1016/j.trgeo.2024.101254
- Apr 18, 2024
- Transportation Geotechnics
Optimized machine learning models for predicting crown convergence of plateau mountain tunnels
- Preprint Article
1
- 10.5194/egusphere-egu2020-4233
- Mar 23, 2020
<p>Accurate water level (WL) forecasting is important for water resources management and planning purposes in the Great Lakes. The objectives of this research are two-fold.  The first objective is to apply machine learning (ML) (i.e., random forest (RF) and support vector regression (SVR)) and hybrid convolutional neural network(CNN)-long-short term memory (LSTM) deep learning (DL) models for multi-step (i.e., one-, two- and three-monthly step ahead) WL forecasting in the Great Lakes (Michigan and Ontario). The second objective is to integrate the boundary corrected (BC) maximal overlap discrete wavelet transform (MODWT) with SVR, RF, and CNN-LSTM models to improve the performance of the individual models. By employing a BC-wavelet decomposition method, the ‘future data’ issue (i.e., data from the future that is not available), often overlooked in the literature and a major barrier to achieving realistic forecasting performance is overcome. </p><p>For Lakes Michigan and Ontario, 1212 monthly WL (m) records (spanning Jan 1918–Dec 2018) were used to develop the models. For the non-wavelet-based models (SVR, RF, and CNN-LSTM), candidate model inputs included the WL recorded over the previous 12 months.  For the BC-MODWT-based models (BC-MODWT-SVR, BC-MODWT-RF, and BC-MODWT-CNN-LSTM), the lagged input time series were decomposed into BC-wavelet and scaling coefficients by using different mother wavelets (Haar, Daubechies, Symlets, Fejer-Korovkin and Coiflets), filter lengths (from two up to 12) and decomposition levels (from one up to seven).  For each method (SVR, RF, and CNN-LSTM), mother wavelet, and decomposition level a model was generated.  For both wavelet- and non-wavelet-based models, the particle swarm optimization (PSO) method was used to select the most appropriate inputs to include in the proposed multi-step WL forecasting models.</p><p>The datasets were partitioned into calibration and validation subsets. After calibrating the models, various performance evaluation metrics, e.g., coefficient of determination (R<sup>2</sup>), root mean square error (RMSE), mean absolute error (MAE), root mean square percentage error (RMSPE), mean absolute percentage error (MAPE) and the Nash-Sutcliffe efficiency coefficient (NSC) were used to assess model accuracy.</p><p>Of the ML models, the SVR outperformed RF while the DL models outperformed the ML models for each forecast lead time (one-, two-, and three-step(s) ahead). Results from this case study indicate that not all wavelet families and decomposition levels perform equally and in some cases, the wavelet-based models do not improve performance over the non-wavelet-based models. However, the BC-MODWT-CNN-LSTM using suitable mother wavelets (e.g., Haar) outperforms the individual ML and BC-MODWT-ML-based models. More accurate forecasts were obtained for Lake Michigan although the performance in both Great Lakes was accurate. The outcomes of this research indicate that the BC-MODWT-CNN-LSTM model is a promising tool for generating accurate WL forecasts.</p>
- Research Article
11
- 10.3390/app12094770
- May 9, 2022
- Applied Sciences
In this study, leaf area prediction models of Dendrobium nobile, were developed through machine learning (ML) techniques including multiple linear regression (MLR), support vector regression (SVR), gradient boosting regression (GBR), and artificial neural networks (ANNs). The best model was tested using the coefficient of determination (R2), mean absolute errors (MAEs), and root mean square errors (RMSEs) and statistically confirmed through average rank (AR). Leaf images were captured through a smartphone and ImageJ was used to calculate the length (L), width (W), and leaf area (LA). Three orders of L, W, and their combinations were taken for model building. Multicollinearity status was checked using Variance Inflation Factor (VIF) and Tolerance (T). A total of 80% of the dataset and the remaining 20% were used for training and validation, respectively. KFold (K = 10) cross-validation checked the model overfit. GBR (R2, MAE and RMSE values ranged at 0.96, (0.82–0.91) and (1.10–1.11) cm2) in the testing phase was the best among the ML models. AR statistically confirms the outperformance of GBR, securing first rank and a frequency of 80% among the top ten ML models. Thus, GBR is the best model imparting its future utilization to estimate leaf area in D. nobile.
- Conference Article
3
- 10.2991/meita-15.2015.118
- Jan 1, 2015
93 compounds which can permeate the placenta barrier were collected as data set for the construction of support vector regression (SVR) model.Besides, 140 compounds with reproductive toxicity and 170 compounds with no reproductive toxicity were collected as another data set for the construction of support vector classification (SVC) model.1481 molecular descriptors were calculated to represent the structure characteristics of all the compounds mentioned above by Dragon2.1.CfsSubsetEval valuation method and BestFirst-D1-N5 searching method were used to optimize the subset of molecular descriptors.Then based on the above data, SVR model for prediction the placenta barrier permeability (PBP) and SVC model for prediction the reproductive toxicity were built respectively by using LibSVM program.Both the SVR model and the SVC model obtained better prediction ability.The correlation coefficient (R 2 ) values of the training set and test set of the optimal SVR model were 0.990 and 0.780.The accuracy, sensitivity, and specificity values of the optimal SVC model were all above 80%.Subsequently, the SVR model was utilized to predict the PBP of the compounds which were collected from 13 commonly used tocolytic Chinese herbs.The compounds with higher permeability were further studied by the SVC model and 15 compounds were classified as positive compounds with reproductive toxicity.The two models constructed in this study might be employed in guiding the application of the tocolytic Chinese herbs in clinical.
- Research Article
115
- 10.1007/s00122-011-1648-y
- Jul 8, 2011
- Theoretical and Applied Genetics
A byproduct of genome-wide association studies is the possibility of carrying out genome-enabled prediction of disease risk or of quantitative traits. This study is concerned with predicting two quantitative traits, milk yield in dairy cattle and grain yield in wheat, using dense molecular markers as predictors. Two support vector regression (SVR) models, ε-SVR and least-squares SVR, were explored and compared to a widely applied linear regression model, the Bayesian Lasso, the latter assuming additive marker effects. Predictive performance was measured using predictive correlation and mean squared error of prediction. Depending on the kernel function chosen, SVR can model either linear or nonlinear relationships between phenotypes and marker genotypes. For milk yield, where phenotypes were estimated breeding values of bulls (a linear combination of the data), SVR with a Gaussian radial basis function (RBF) kernel had a slightly better performance than with a linear kernel, and was similar to the Bayesian Lasso. For the wheat data, where phenotype was raw grain yield, the RBF kernel provided clear advantages over the linear kernel, e.g., a 17.5% increase in correlation when using the ε-SVR. SVR with a RBF kernel also compared favorably to the Bayesian Lasso in this case. It is concluded that a nonlinear RBF kernel may be an optimal choice for SVR, especially when phenotypes to be predicted have a nonlinear dependency on genotypes, as it might have been the case in the wheat data.