Stress release coefficient prediction of sandy-gravel soil by extra tree algorithms
Stress release coefficient prediction of sandy-gravel soil by extra tree algorithms
- Research Article
1
- 10.3390/foods14234069
- Nov 27, 2025
- Foods
This study examines the effectiveness of machine learning approaches for the automatic identification of watermelon genotypes from the seeds of watermelon, for the snack-type watermelon (Citrullus lanatus). Nine genotypes with red, white, and black seed coats were assessed in total. For each genotype, 200 seeds were analyzed using high-resolution imaging and digital measurement techniques for the extraction of morphological characteristics (length, width, thickness, area, perimeter, equivalent diameter, etc., and physical (weight) and colorimetric attributes of the (L, a, b). The resulting dataset was modeled using Artificial Neural Network (ANN), Random Forest (RF) and Extra Tree (ET) algorithms and performance was validated by a 10-fold cross-validation. The primary objective of the study was to match (identify) each seed accurately with its respective genotype by using the morphological, physical, and colorimetric characteristics of the seed and thus to perform genotypic classification. The comparative results showed that the RF model had the highest genotypic performance (accuracy 92.22%, F1-score 91.87%, Cohen’s Kappa 0.9118), followed by the ET (accuracy, 90.00%) and ANN models with a relatively lower precision (86.11%). Statistical analysis using the Wilcoxon signed-rank test confirmed that both RF and ET significantly outperformed ANN, with RF providing superior balance and stability over ET. The findings highlight that machine learning-based frameworks enable rapid, reliable, and non-destructive classification (identification) of snack-type watermelon seeds according to their genotypes. Such approaches hold strong potential for enhancing varietal traceability in breeding programs, improving quality control in commercial seed production, and meeting the high-throughput demands of seed processing industries.
- Conference Article
4
- 10.2118/217116-ms
- Jul 30, 2023
The worst-case discharge during a blowout is a major concern for the oil and gas industry. Various two-phase flow patterns are established in the wellbore during a blowout incident. One of the challenges for field engineers is accurately predicting the flow pattern and, subsequently, the pressure drop along the wellbore to successfully control the well. Existing machine learning models rely on instantaneous pressure drop and liquid hold-up measurements that are not readily available in the field. This study aims to develop a novel machine-learning model to predict two-phase flow patterns in the wellbore for a wide range of inclination angles (0 − 90 degrees) and superficial gas velocities. The model also helps identify the most crucial wellbore parameter that affects the flow pattern of a two-phase flow. This study collected nearly 5000 data points with various flow pattern observations as a data bank for model formulation. The input data includes pipe diameter, gas velocity, liquid velocity, inclination angle, liquid viscosity and density, and visualized/observed flow patterns. As a first step, the observed flow patterns from different sources are displayed in well-established flow regime maps for vertical and horizontal pipes. The data set was graphically plotted in the form of a scatter matrix, followed by statistical analysis to eliminate outliers. A number of machine learning algorithms are considered to develop an accurate model. These include Support Vector Machine (SVM), Multi-layer Perceptron (MLP), Gradient Boosting algorithm, CatBoost, and Extra Tree algorithm, and the Random Forest algorithm. The predictive abilities of the models are cross compared. Because of their unique features, such as variable-importance plots, the CatBoost, Extra Tree, and Random Forest algorithms are selected and implemented in the model to determine the most crucial wellbore parameters affecting the two-phase flow pattern. The Variable-importance plot feature makes CatBoost, Extra Tree, and Random Forest the best option for investigating two-phase flow characteristics using machine learning techniques. The result showed that the CatBoost model predictions demonstrate 98% accuracy compared to measurements. Furthermore, its forecast suggests that in-situ superficial gas velocity is the most influential variable affecting flow pattern, followed by superficial liquid velocity, inclination angle, pipe diameter, and liquid viscosity. These findings could not be possible with the commonly used empirical correlations. For instance, according to previous phenomenological models, the impact of the inclination angle on the flow pattern variation is negligible at high in-situ superficial gas velocities, which contradicts the current observation. The new model requires readily available field operating parameters to predict flow patterns in the wellbore accurately. A precise forecast of flow patterns leads to accurate pressure loss calculations and worst-case discharge predictions.
- Research Article
1
- 10.5747/ca.2024.v20.h526
- Apr 15, 2024
- COLLOQUIUM AGRARIAE
Cotton has a considerable economic impact on agribusiness. Strategies to reduce production loss due, for example, to pest attacks are increasingly required. Spodoptera frugiperda, known as fall armyworm, causes irreversible damage to cotton. In this context, a current approach is the use of hyperspectral measurements obtained by remote sensors and processed by machine learning algorithms. However, such measures generate data redundancy, making it difficult to extract information. An alternative is to apply pre-processing techniques, but little is known about the impact these generate on the learning ability of algorithms. This study evaluates the performance of machine learning algorithms in identifying cotton plants attacked by pests using pre-processed and raw hyperspectral measurements. Data are collected by EMBRAPA, and consist of hyperspectral measurements, in the range of 350-2500 nm, referring to eight days of collections in healthy cotton plants and attacked by S. frugiperda. Pre-processing techniques to try are baseline removal, smoothing, first and second order derivatives. A group of machine learning algorithms, such as Random Forest, Support Vector Machine, Extra Tree, was used to model pre-processed and non-pre-processed hyperspectral measurements. According to the proposed metric, the F-Score and the Extra Trees (ExT) algorithm performed better (0.77). So it overlapped the other results with the preprocessed dataset. In addition to obtaining the most important lengths for the algorithm to have its best performance. Concluding that machine learning with spectroscopy can help the field in a promising way. Studies in other crops and with other factors applied to the plant are recommended.
- Research Article
52
- 10.1016/j.petlm.2021.03.001
- Mar 5, 2021
- Petroleum
Application of artificial intelligence in predicting the dynamics of bottom hole pressure for under-balanced drilling: Extra tree compared with feed forward neural network model
- Research Article
36
- 10.3390/rs15174214
- Aug 27, 2023
- Remote Sensing
Soil moisture is a key parameter for the circulation of water and energy exchange between surface and the atmosphere, playing an important role in hydrology, agriculture, and meteorology. Traditional methods for monitoring soil moisture suffer from spatial discontinuity, time-consuming processes, and high costs. Remote sensing technology enables the non-destructive and efficient retrieval of land information, allowing rapid soil moisture monitoring to schedule crop irrigation and evaluate the irrigation efficiency. Satellite data with different resolutions provide different observation scales. Evaluating the accuracy of estimating soil moisture based on open and free satellite data, as well as exploring the comprehensiveness and adaptability of different satellites for soil moisture temporal and spatial observations, are important research contents of current soil moisture monitoring. The study utilized three types of satellite data, namely GF-1, Landsat-8, and GF-4, with respective temporal and spatial resolutions of 16 m (every 4 days), 30 m (every 16 days), and 50 m (daily). The gray relational analysis (GRA) was employed to identify vegetation indices that selected sensitivity to soil moisture at varying depths (3 cm, 10 cm, and 20 cm). Then, this study employed random forest (RF), Extra Tree (ETr), and linear regression (LR) algorithms to estimate soil moisture at different depths with optical satellite data sources. The results showed that the accuracy of soil moisture estimation was different at different growth stages. The model accuracy exhibited an upward trend during the middle and late growth stages, coinciding with higher vegetation coverage; however, it demonstrated a decline in accuracy during the early and late growth stages due to either the absence or limited presence of vegetation. Among the three satellite images, the vegetation indices derived from GF-1 exhibited were more sensitive to vegetation characteristics and demonstrated superior soil moisture estimation accuracy (with R2 ranging 0.129–0.928, RMSE ranging 0.017–0.078), followed by Landsat-8 (with R2 ranging 0.117–0.862, RMSE ranging 0.017–0.088). The soil moisture estimation accuracy of GF-4 was the worst (with R2 ranging 0.070–0.921, RMSE ranging 0.020–0.140). Thus, GF-1 is suitable for vegetated areas. In addition, the ETr model outperformed the other models in both accuracy and stability (ETr model: R2 ranging from 0.117 to 0.928, RMSE ranging from 0.021 to 0.091; RF model: R2 ranging from 0.225 to 0.926, RMSE ranging from 0.019 to 0.085; LR model: R2 ranging from 0.048 to 0.733, RMSE ranging from 0.030 to 0.144). Utilizing GF-1 is recommended to construct the ETr model for assessing soil moisture variations in the farming land of northern China. Therefore, in cases where there are limited ground sample data, it is advisable to utilize high-spatiotemporal-resolution remote sensing data, along with machine learning algorithms such as ETr and RF, which are suitable for small samples, for soil moisture estimation.
- Research Article
31
- 10.1109/tte.2021.3127194
- Jun 1, 2022
- IEEE Transactions on Transportation Electrification
In vehicle operation, in order to maximize the fuel economy, the propulsion system control can easily adapt to pressure and temperature variations as these variations can be measured by sensors. However, it is challenging to detect driving cycles. With growing progress made in the artificial intelligence field, pattern recognition gains momentum in various applications. This study presents a study on driving cycle pattern recognition based on supervised learning. Training data 2-D visualization is achieved by the t-distributed stochastic neighbor embedding (t-SNE) algorithm. Ten out of 12 supervised learning algorithms predict driving conditions with an accuracy of 88% or higher, and the extra tree (ET) algorithm leads the recognition accuracy at 90.26%. To improve the recognition accuracy, two hierarchical frameworks are proposed by integrating multiple supervised learning methods using weighted average and vote methods. The two hierarchical frameworks boost the driving condition recognition accuracy from 90.26% to 90.43% (weighted average) and 91.76% (vote). The results are further validated in the holdout test. In addition, a plug-in hybrid electric vehicle simulation shows 3.88%–5.82% fuel economy improvement compared to the baseline method. Two hierarchical methods outperform the ET method by 2% fuel economy. In summary, supervised learning shows great potential to detect driving cycles for vehicle energy saving.
- Research Article
2
- 10.1186/s43046-025-00265-3
- Mar 17, 2025
- Journal of the Egyptian National Cancer Institute
BackgroundMachine learning (ML) is a significant area of artificial intelligence, which can improve the accuracy of predictive or diagnostic models for differentiating between prostate biopsy outcomes. This study aims to develop a novel decision-support ML model for classifying patients with biopsy-negative (cancer-free), clinically significant, and non-clinically significant prostate cancer across two prostate-specific antigen (PSA) cut-offs ≤ 10 ng/ml and > 10 ng/ml.MethodsThe data for the current study were retrieved from the records of two main hospitals in Riyadh, Saudi Arabia from July 2018 through July 2024. Six machine learning algorithms were employed, and the dataset was randomly divided into a training set and a validation set at a ratio of 8:2. The following metrics were used as performance indicators across the six algorithms: Accuracy, Precision, Recall, F1-score, and area under the curve. Recent data from the two hospitals was utilized for external validation.ResultsThe metrics for Random Forest, Extra Tree, and Decision Tree algorithms showed excellent capability in classifying the outcomes of prostate biopsy for the two PSA cut-offs. However, the metrics for the PSA cut-off > 10 ng/ml are higher than those for PSA ≤ 10 ng/ml. For the three-class classification, the accuracy and area under the curve for the cut-off > 10 ng/ml were 0.96 and 0.99, respectively. While for the cut-off ≤ 10 ng/ml they were 0.92 and 0.94 for Random Forest and 0.94 and 0.95 for the Extra Tree algorithm. The metrics of non-clinically significant and biopsy-negative cases outperformed those of clinically significant cases.ConclusionML models are proving to be effective tools in differentiating between prostate biopsy outcomes, enhancing diagnostic accuracy, and potentially transforming clinical practices in prostate cancer management.
- Research Article
6
- 10.4108/eetiot.2269
- Nov 30, 2023
- EAI Endorsed Transactions on Internet of Things
The search for effective solutions to address traffic congestion presents a significant challenge for large urban cities. Analysis of urban traffic congestion has revealed that more than 70% of it can be attributed to prolonged searches for parking spaces. Consequently, accurate prediction of parking space availability in advance can play a vital role in assisting drivers to find vacant parking spaces quickly. Such solutions hold the potential to reduce traffic congestion and mitigate its detrimental impacts on the environment, economy, and public health. Machine learning algorithms have emerged as promising approaches for predicting parking space availability. However, comparative studies on those machine learning models to evaluate the best suited for a large-scale prediction and within a given prediction time period are missing.In this study, we compared nine machine learning algorithms to assess their efficiency in predicting long-term, large-scale parking space availability. Our comparison was based on two approaches: using on-street parking data alone and 2) incorporating data from external sources (such as weather data). We used automatic machine learning models to compare the performance of different algorithms according to the prediction efficiency and execution time. Our results indicated that the automated machine learning models implemented were well fitted to our data. Notably, the Extra Tree and Random Forest algorithms demonstrated the highest efficiency among the models tested. Moreover, we observed that the Random Forest algorithm exhibited less computational demand than the Extra Tree algorithm, making it particularly advantageous in terms of execution time. Therefore, this work suggests that the Random Forest algorithm is the most suitable machine learning model in terms of efficiency and execution time for accurately predicting large-scale, long-term parking space availability.
- Research Article
- 10.59395/ijadis.v6i3.1428
- Dec 1, 2025
- International Journal of Advances in Data and Information Systems
Gastroesophageal reflux disease (GERD) is a prevalent gastrointestinal disorder characterized by the backward flow of gastric contents into the esophagus, often causing heartburn and regurgitation, with a global prevalence of approximately 13.98%. Early detection is essential to prevent severe complications such as esophagitis, esophageal strictures, and esophageal cancer. However, conventional diagnostic methods are often limited by inadequate healthcare resources and high cost, particularly in developing countries. On the other hand, machine learning can be implemented as a promising alternative method for disease detection, improving accuracy through data pattern identification. Machine learning has been used for several disease detection tasks, such as Breast Cancer, Diabetes, etc. This study proposed an enhanced GERD prediction model by implementing the Extra Tree classifier optimized by the Komodo Mlipir Algorithm (KMA) for hyperparameter optimization. This study used a GERD dataset from the Harvard Dataverse, which consists of 1200 rows with 69 features. The result shows that the Extra Tree Algorithm that KMA tuned achieved a high-performance evaluation with an F1-score of 0.97. This highlights the effectiveness of KMA in enhancing model performance. Compared to the previous study, the proposed Extra Tree Models optimized by KMA performed improved performance, demonstrating the effectiveness of metaheuristic optimization in GERD prediction.
- Research Article
22
- 10.1016/j.ijbiomac.2022.09.202
- Sep 25, 2022
- International Journal of Biological Macromolecules
Development and analysis of machine-learning guided flash nanoprecipitation (FNP) for continuous chitosan nanoparticles production
- Research Article
103
- 10.3390/biomedicines11020581
- Feb 16, 2023
- Biomedicines
There has been a sharp increase in liver disease globally, and many people are dying without even knowing that they have it. As a result of its limited symptoms, it is extremely difficult to detect liver disease until the very last stage. In the event of early detection, patients can begin treatment earlier, thereby saving their lives. It has become increasingly popular to use ensemble learning algorithms since they perform better than traditional machine learning algorithms. In this context, this paper proposes a novel architecture based on ensemble learning and enhanced preprocessing to predict liver disease using the Indian Liver Patient Dataset (ILPD). Six ensemble learning algorithms are applied to the ILPD, and their results are compared to those obtained with existing studies. The proposed model uses several data preprocessing methods, such as data balancing, feature scaling, and feature selection, to improve the accuracy with appropriate imputations. Multivariate imputation is applied to fill in missing values. On skewed columns, log1p transformation was applied, along with standardization, min–max scaling, maximum absolute scaling, and robust scaling techniques. The selection of features is carried out based on several methods including univariate selection, feature importance, and correlation matrix. These enhanced preprocessed data are trained on Gradient boosting, XGBoost, Bagging, Random Forest, Extra Tree, and Stacking ensemble learning algorithms. The results of the six models were compared with each other, as well as with the models used in other research works. The proposed model using extra tree classifier and random forest, outperformed the other methods with the highest testing accuracy of 91.82% and 86.06%, respectively, portraying our method as a real-world solution for detecting liver disease.
- Research Article
47
- 10.1177/01445987221138135
- Nov 14, 2022
- Energy Exploration & Exploitation
Recently, power systems have faced the challenges of growing electricity demand, reducing fossil fuels, and exacerbating environmental pollution due to carbon emissions from fossil fuel-based power generation. Integrating low-carbon alternative energy, renewable energy sources (RES), is becoming very important for energy systems. Effective management of the integration of the production capacity of RES is as important as the production capacity of wind farms with the production capacity of fossil fuel power plants. This article analyzed 850,660 data recorded by a wind farm from March 01, 2020, 00:00:00 to December 31, t2020, 23:50:00 were analyzed. And by using machine learning and extra tree, light gradient boosting machine, gradient boosting regressor, decision tree, Ada Boost, and ridge algorithms, the production power of the wind farm was predicted. The best performance predicting the turbine production power was assigned to extra tree, and the worst performance was related to the Ridge algorithm.
- Research Article
142
- 10.1007/s10772-021-09837-9
- Apr 14, 2021
- International Journal of Speech Technology
Parkinson’s disease is a neurodegenerative disorder that progresses slowly and its symptoms appear over time, so its early diagnosis is not easy. A neurologist can diagnose Parkinson's by reviewing the patient's medical history and repeated scans. Besides, body movement analysts can diagnose Parkinson's by analyzing body movement. Recent research work has shown that changes in speech can be used as a measurable indicator for early Parkinson’s detection. In this work, the authors propose a speech signal-based hybrid Parkinson's disease diagnosis system for its early diagnosis. To achieve this, the authors have tested several combinations of feature selection approaches and classification algorithms and designed the model with the best combination. To formulate various combinations, three feature selection methods such as mutual information gain, extra tree, and genetic algorithm and three classifiers namely naive bayes, k-nearest-neighbors, and random forest have been used. To analyze the performance of different combinations, the speech dataset available at the UCI (University of California, Irvine) machine learning repository has been used. As the dataset is highly imbalanced so the class balancing problem is overcome by the synthetic minority oversampling technique (SMOTE). The combination of genetic algorithm and random forest classifier has shown the best performance with 95.58% accuracy. Moreover, this result is also better than the recent work found in the literature.
- Research Article
17
- 10.4103/ijo.ijo_2989_22
- May 1, 2023
- Indian Journal of Ophthalmology
Purpose:Recently, the proportion of patients with high myopia has shown a continuous growing trend, more toward the younger age groups. This study aimed to predict the changes in spherical equivalent refraction (SER) and axial length (AL) in children using machine learning methods.Methods:This study is a retrospective study. The cooperative ophthalmology hospital of this study collected data on 179 sets of childhood myopia examinations. The data collected included AL and SER from grades 1 to 6. This study used the six machine learning models to predict AL and SER based on the data. Six evaluation indicators were used to evaluate the prediction results of the models.Results:For predicting SER in grade 6, grade 5, grade 4, grade 3, and grade 2, the best results were obtained through the multilayer perceptron (MLP) algorithm, MLP algorithm, orthogonal matching pursuit (OMP) algorithm, OMP algorithm, and OMP algorithm, respectively. The R2 of the five models were 0.8997, 0.7839, 0.7177, 0.5118, and 0.1758, respectively. For predicting AL in grade 6, grade 5, grade 4, grade 3, and grade 2, the best results were obtained through the Extra Tree (ET) algorithm, MLP algorithm, kernel ridge (KR) algorithm, KR algorithm, and MLP algorithm, respectively. The R2 of the five models were 0.7546, 0.5456, 0.8755, 0.9072, and 0.8534, respectively.Conclusion:Therefore, in predicting SER, the OMP model performed better than the other models in most experiments. In predicting AL, the KR and MLP models were better than the other models in most experiments.
- Research Article
37
- 10.3390/plants11131697
- Jun 27, 2022
- Plants (Basel, Switzerland)
Current development in precision agriculture has underscored the role of machine learning in crop yield prediction. Machine learning algorithms are capable of learning linear and nonlinear patterns in complex agro-meteorological data. However, the application of machine learning methods for predictive analysis is lacking in the oil palm industry. This work evaluated a supervised machine learning approach to develop an explainable and reusable oil palm yield prediction workflow. The input data included 12 weather and three soil moisture parameters along with 420 months of actual yield records of the study site. Multisource data and conventional machine learning techniques were coupled with an automated model selection process. The performance of two top regression models, namely Extra Tree and AdaBoost was evaluated using six statistical evaluation metrics. The prediction was followed by data preprocessing and feature selection. Selected regression models were compared with Random Forest, Gradient Boosting, Decision Tree, and other non-tree algorithms to prove the R2 driven performance superiority of tree-based ensemble models. In addition, the learning process of the models was examined using model-based feature importance, learning curve, validation curve, residual analysis, and prediction error. Results indicated that rainfall frequency, root-zone soil moisture, and temperature could make a significant impact on oil palm yield. Most influential features that contributed to the prediction process are rainfall, cloud amount, number of rain days, wind speed, and root zone soil wetness. It is concluded that the means of machine learning have great potential for the application to predict oil palm yield using weather and soil moisture data.