Tree-based machine learning methods for predicting vehicle insurance claim size
This study compares classical parametric models and tree-based ensemble methods for predicting vehicle insurance claim size, finding that ensemble methods offer modest accuracy improvements over traditional models, with key predictors identified as premium and insured value, while the Tweedie GLM remains a competitive benchmark.
Vehicle insurance claim severity modeling requires accurate and interpretable methods that can handle skewed and heterogeneous loss data. This study provides a structured empirical comparison between classical parametric regression models and tree-based ensemble learning approaches for predicting claim size conditional on claim occurrence. The analysis is conducted within a cross-sectional conditional severity framework using real-world motor insurance data. We implement and compare ordinary least squares (OLS), a Tweedie generalized linear model (GLM), and three ensemble methods: bagging, random forests (RFs), and gradient boosting. Model performance is evaluated using out-of-sample root mean square error (RMSE), and variable importance measures assess the relative contribution of predictors. The results indicate that tree-based ensemble methods achieve modest improvements in predictive accuracy relative to classical parametric models. The Tweedie GLM remains a competitive, flexible parametric benchmark for skewed positive claim amounts. Variable importance analysis consistently identifies premium and insured value as key determinants of claim severity. Overall, the findings suggest that ensemble learning methods can complement traditional actuarial models, offering additional flexibility in capturing non-linear effects while maintaining comparable predictive performance in moderate-complexity severity data.
- Research Article
34
- 10.1016/j.asoc.2023.110041
- Jan 20, 2023
- Applied Soft Computing
A Health state-related ensemble deep learning method for aircraft engine remaining useful life prediction
- Research Article
392
- 10.1007/s42452-020-3060-1
- Jun 30, 2020
- SN Applied Sciences
Decision tree-based classifier ensemble methods are a machine learning (ML) technique that combines several tree models to produce an effective or optimum predictive model, and that allows well-predictive performance especially compared to a single model. Thus, selecting a proper ML algorithm help us to understand possible future occurrences by analyzing the past more accurate. The main purpose of this study is to produce landslide susceptibility map of the Ayancik district of Sinop province, situated in the Black Sea region of Turkey using three featured regression tree-based ensemble methods including gradient boosting machines (GBM), extreme gradient boosting (XGBoost), and random forest (RF). Fifteen landslide causative factors and 105 landslide locations occurred in the region were used. The landslide inventory map was randomly divided into training (70%) and testing (30%) dataset to construct the RF, XGBoost and GBM prediction models. Symmetrical uncertainty measure was utilized to determine the most important causative factors, and then the selected features were used to construct susceptibility prediction models. The performance of the ensemble models was validated using different accuracy metrics including Area under the curve (AUC), overall accuracy (OA), Root mean square error (RMSE), and Kappa coefficient. Also, the Wilcoxon signed-rank test was used to assess differences between optimum models. The accuracy results showed that the model of XgBoost_Opt model (the model created by optimum factor combination) has the highest prediction capability (OA = 0.8501 and AUC = 0.8976), followed by the RF_opt (OA = 0.8336 and AUC = 0.8860) and GBM_Opt (OA = 0.8244 and AUC = 0.8796). When the Wilcoxon sign-rank test results were analyzed, XgBoost_Opt model, which is the best subset combinations, were confirmed to be statistically significant considering other models. The results showed that, the XGBoost method according to optimum model achieved lower prediction error and higher accuracy results than the other ensemble methods.
- Research Article
44
- 10.1016/j.apr.2021.03.008
- Mar 30, 2021
- Atmospheric Pollution Research
Ensemble multifeatured deep learning models for air quality forecasting
- Research Article
6
- 10.1155/2022/2590940
- Oct 25, 2022
- Mathematical Problems in Engineering
Option pricing based on data-driven methods is a challenging task that has attracted much attention recently. There are mainly two types of methods that have been widely used, respectively, the neural network method and the ensemble learning method. The option pricing model based on the neural network has high complexity, and a large number of hyper-parameters will be generated during training, resulting in difficult model adjustment. Furthermore, a lot of training data are needed. The option pricing model based on ensemble learning is not ideal for data feature extraction, because each calculation of the ensemble learning method is mainly to reduce the final residual. Therefore, this paper adopts a learning framework that embeds the modular ensemble learning methods into the network learning structure, and an option pricing model based on deep ensemble learning is proposed. The model is mainly composed of two parts: features reorganization based on random forest, used to calculate the importance of features, combined with the original data as training input; the multilayer ensemble data training structure is based on network learning structure and embeds two ensemble learning methods as network modules, and it also designs a stop algorithm to automatically determine the number of layers. This enables the model to retain the effect of data feature extraction and adapt to small and medium data sets without generating many hyper-parameters. Moreover, in order to make the model fully absorb the advantages of the two ensemble learning methods, we adopt cross-training for data. From the experimental results, it can be concluded that compared with the current optimal method, the prediction performance of the proposed model is improved by 36% in the root mean square error (RMSE), which proves the superiority of the proposed model from the quantitative direction.
- Research Article
18
- 10.7717/peerj-cs.2234
- Aug 7, 2024
- PeerJ. Computer science
The continuous increase in carbon dioxide (CO2) emissions from fuel vehicles generates a greenhouse effect in the atmosphere, which has a negative impact on global warming and climate change and raises serious concerns about environmental sustainability. Therefore, research on estimating and reducing vehicle CO2 emissions is crucial in promoting environmental sustainability and reducing greenhouse gas emissions in the atmosphere. This study performed a comparative regression analysis using 18 different regression algorithms based on machine learning, ensemble learning, and deep learning paradigms to evaluate and predict CO2 emissions from fuel vehicles. The performance of each algorithm was evaluated using metrics including R2, Adjusted R2, root mean square error (RMSE), and runtime. The findings revealed that ensemble learning methods have higher prediction accuracy and lower error rates. Ensemble learning algorithms that included Extreme Gradient Boosting (XGB), Random Forest, and Light Gradient-Boosting Machine (LGBM) demonstrated high R2 and low RMSE values. As a result, these ensemble learning-based algorithms were discovered to be the most effective methods of predicting CO2 emissions. Although deep learning models with complex structures, such as the convolutional neural network (CNN), deep neural network (DNN) and gated recurrent unit (GRU), achieved high R2 values, it was discovered that they take longer to train and require more computational resources. The methodology and findings of our research provide a number of important implications for the different stakeholders striving for environmental sustainability and an ecological world.
- Research Article
46
- 10.1109/access.2022.3163291
- Jan 1, 2022
- IEEE Access
The ability to predict the radioactive soil radon gas concentration is important for human beings because it serves as a precursor to earthquakes. Several studies have been conducted across the globe to confirm the correlation of radon emission dynamics and earthquakes, and concluded that the soil radon gas is the witness of anomalous behaviour before the occurrences of several earthquakes. This anomalous behavior can help to construct a better prediction model for earthquake forecasting. This paper aims at employing different ensemble and individual machine learning methods on real time radon time series data with different scenarios to predict anomalies in data caused by the seismic activities.The ensemble methods include boosted tree, bagged cart and boosted linear model while standalone machine learning methods include support vector machine with linear and radial kernels and k-nearest neighbors ( <inline-formula xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink"> <tex-math notation="LaTeX">${K}$ </tex-math></inline-formula> -NN). We tested the methods on a dataset recorded on the fault line located in Muzaffarabad. Time series data was collected over a period ranging from March 1, 2017 to May 11, 2018 including nine(09) earthquakes. The methods are tested in four different settings with 10 times 10 folds cross validation procedure over the time window of 1 to 4. The repeated 10 fold cross validation is performed to reduce the noise in the model performance estimation by replicating the 10 fold cross validation procedure 10 times. Statistical performance evaluation measures viz. root mean square error (RMSE), root mean squared log error (RMSLE), mean absolute percentage error (MAPE), percentage bias (PB), and mean squared error (MSE) have been calculated for the assessment of performance. In setting 1, the support vector machine with radial kernel performs better with the minimum RMSE score of 1381.023 when compared to other prediction models. In setting 3, it can be observed through different performance metrics such as RMSE, the value in the range [1262.864, 1409.616] which is minimum when other prediction models for predicting soil radon gas concentration dataset. For setting 4, the boosted tree model yielded the minimum RMSE and MAPE scores of 1573.174 and 0.056 respectively. Findings of the study shows that boosted tree and support vector machine with radial kernel proved to be better regression models for the prediction of anomalies in soil radon gas concentration during seismic activities. An important finding of this study suggests that by employing boosted tree ensemble method make us able to accurately predict soil radon gas concentration automatically from environmental parameters.
- Research Article
54
- 10.1016/j.istruc.2022.10.056
- Oct 25, 2022
- Structures
An interpretable ensemble learning method to predict the compressive strength of concrete
- Research Article
5
- 10.19127/mbsjohs.889492
- Apr 30, 2021
- Middle Black Sea Journal of Health Science
Objective: In recent years, ensemble learning methods have gained widespread use for early diagnosis of cancer diseases. In this study, it is aimed to establish a high-performance ensemble learning model for early diagnosis and classification of renal cell carcinomas.Methods: In the study, hemogram and laboratory data of 140 patients with renal cell carcinoma and 140 patients without renal cell carcinoma were included in the study. The data set includes 27 predictors and 1 dependent variable. The data were obtained retrospectively. In the study, classification performances of machine learning methods and ensemble learning methods were compared. In the study, classification performances of boosting, bagging, voting and stacking ensemble learning methods as well as IB1, IBk, Kstar, LWL, REPTree, Random Forest and SMO classifiers were compared.Results: REPTree classifier provided the highest performance among machine learning methods (Accuracy = 0.867). Among the ensemble learning methods, the Stacking ensemble learning method provided the highest performance in Model 6 (Accuracy = 0.906). Stacking ensemble learning methods performed higher than boosting, voting, bagging ensemble methods and machine learning methods.Conclusion: Stacking ensemble learning methods provide successful results in the early diagnosis of renal cell carcinomas. Stacking ensemble learning methods can be used as an alternative to existing methods for diagnosing renal cell carcinoma. In order to further increase the classification performance of the stacking ensemble learning method, it is recommended to choose a meta classifier suitable for the data set and variable types.
- Research Article
1
- 10.1504/ijmissp.2017.088165
- Jan 1, 2017
- International Journal of Machine Intelligence and Sensory Signal Processing
Ensemble methods such as boosting combine multiple learners to obtain better prediction than could be obtained from any individual learner. Here we propose a principled framework for directly constructing ensemble learning methods from kernel methods. Unlike previous studies showing the equivalence between boosting and support vector machines (SVMs) which need a translation procedure, we show that it is possible to design boosting-like procedure to solve the SVM optimisation problems. In other words, it is possible to design ensemble methods directly from SVM without any middle procedure. This finding not only enables us to design new ensemble learning methods directly from kernel methods, but also makes it possible to take advantage of those highly-optimised fast linear SVM solvers for ensemble learning. The resulted model is as effective as kernel methods while being as efficient as ensemble methods. We exemplify this framework for designing new binary and multi-class classification ensemble learning as well as a new quantile regression ensemble learning method. Experimental results demonstrate the flexibility and usefulness of the proposed framework.
- Research Article
2
- 10.1504/ijmissp.2017.10009116
- Jan 1, 2017
- International Journal of Machine Intelligence and Sensory Signal Processing
Ensemble methods such as boosting combine multiple learners to obtain better prediction than could be obtained from any individual learner. Here we propose a principled framework for directly constructing ensemble learning methods from kernel methods. Unlike previous studies showing the equivalence between boosting and support vector machines (SVMs) which need a translation procedure, we show that it is possible to design boosting-like procedure to solve the SVM optimisation problems. In other words, it is possible to design ensemble methods directly from SVM without any middle procedure. This finding not only enables us to design new ensemble learning methods directly from kernel methods, but also makes it possible to take advantage of those highly-optimised fast linear SVM solvers for ensemble learning. The resulted model is as effective as kernel methods while being as efficient as ensemble methods. We exemplify this framework for designing new binary and multi-class classification ensemble learning as well as a new quantile regression ensemble learning method. Experimental results demonstrate the flexibility and usefulness of the proposed framework.
- Research Article
10
- 10.31127/tuje.1352481
- Apr 30, 2024
- Turkish Journal of Engineering
The coronavirus pandemic has distanced people from social life and increased the use of social media. People's emotions can be determined with text data collected from social media applications. This is used in many fields, especially in commerce. This study aims to predict people's sentiments about the pandemic by applying sentiment analysis to Twitter tweets about the pandemic using single machine learning classifiers (Decision Tree-DT, K-Nearest Neighbor-KNN, Logistic Regression-LR, Naïve Bayes-NB, Random Forest-RF) and ensemble learning methods (Majority Voting (MV), Probabilistic Voting (PV), and Stacking (STCK)). After vectorizing the tweets using two predictive methods, Word2Vec (W2V) and Doc2Vec, and two traditional word representation methods, Term Frequency-Inverse Document Frequency (TF-IDF) and Bag of Words (BOW), classification models built using single machine learning classifiers were compared to models built using ensemble learning methods (MV, PV and STCK) by heterogeneously combining single machine classifier algorithms. Accuracy (ACC), F-measure (F), precision (P), and recall (R) were used as performance measures, with training/test separation rates of 70%-30% and 80%-20%, respectively. Among these models, the ACC of ensemble learning models ranged from 89% to 73%, while the ACC of single classifier models ranged from 60% to 80%. Among the ensemble learning methods, STCK with Doc2Vec text representation/embedding method gave the best ACC result of 89%. According to the experimental results, ensemble models built with heterogeneous machine learning classifier algorithms gave better results than single machine learning classifier algorithms.
- Research Article
2
- 10.1177/13694332251327844
- Mar 19, 2025
- Advances in Structural Engineering
Local scour is one of the main reasons for bridge collapse. To solve the difficult problem of detecting the local scour depth of underwater pier structures, this paper explores an optimal method for predicting the local scour depth of underwater pier structures based on various ensemble learning methods. Firstly, this paper collects 487 sets of data samples containing nine input parameters with corresponding scour depths from the open-source database in the practical project. Secondly, this paper employs five algorithms commonly used in ensemble learning, that is, Random Forest (RF), Gradient Boosted Decision Tree (GBDT), Extreme Gradient Boosting (XGBoost), Adaptive Boosting (AdaBoost), and Light Gradient Boosting Machine (LightGBM), to build a prediction model of the local scour depth. In addition, the Bayesian hyperparameter optimization method is applied to search for the best hyperparameter combination of the model. Then, eight evaluation indices, including Mean Absolute Error (MAE), Mean Bias Error (MBE), Mean Absolute Percentage Error (MAPE), Root Mean Square Error (RMSE), coefficient of determination (R 2 ), Nash-Sutcliffe Efficiency (NSE), Percent Bias (Pbias), and Willmott Index (WI), were used to compare and analyse the established prediction model, and the importance coefficients of each input parameter were evaluated based on this prediction model. Finally, Conditional Generative Adversarial Network (CGAN) was applied to augment and supplement the samples in the existing database, and the prediction model was used to verify its effectiveness. The results of this paper show that the parameter-optimized LightGBM model achieves the best prediction performance. Moreover, the established CGAN model can effectively solve the problem of insufficient data samples and lack of specific sample data.
- Dissertation
2
- 10.32657/10220/46281
- Jan 1, 2018
In recent years, time series forecasting has obtained significant academic and industrial interest with its significance in various application fields, including power system related applications (electric load, wind power and solar irradiance forecasting, etc.), as well as financial market related applications (stock price, exchange rate and electricity price forecasting, etc.). Many statistics based machine learning models have been proposed to obtain accurate results for time series forecasting in the literature. The methods can be divided into two categories: linear models (such as auto-regressive moving average) and non-linear models (such as artificial neural network and support vector machine). However, due to the highly nonlinear characteristics of real world time series signals caused by various influencing factors, it is very difficult to ensure the performance of machine learning models. Deep learning and ensemble methods are possible solutions to this problem. This thesis mainly focuses on the state-of-the-art ensemble learning methods and deep learning models for both power system and financial market related time series forecasting. The development of time series forecasting is introduced, and a brief review of existing algorithms is also recorded. Motivated by the attractive advantages of ensemble learning, two deep learning based ensemble methods are presented: (i) ensemble method composed of deep belief networks and support vector machines, (ii) empirical mode decomposition (EMD) based ensemble deep learning model. The performance of the proposed methods is evaluated by real world time series datasets. On the other hand, the ensemble methods based on fast learning models are also investigated in this thesis, such as decision tree ensembles and random vector functional link (RVFL) network based hybrid models. Specifically, a novel decomposition method composed of discrete wavelet transform and EMD is combined with incremental RVFL for electric load forecasting. Finally, the advantages and potential future developments of deep learning and ensemble methods are discussed.
- Research Article
375
- 10.1016/j.dss.2013.08.002
- Aug 15, 2013
- Decision Support Systems
Sentiment classification: The contribution of ensemble learning
- Research Article
53
- 10.1016/j.egyr.2022.10.298
- Oct 28, 2022
- Energy Reports
RUL Prediction for Lithium Batteries Using a Novel Ensemble Learning Method