Accelerate Literature Icon
Want to do a literature review? Try our new Literature Review workflow

The usability of stacking-based ensemble learning model in crime prediction: a systematic review

  • Abstract
  • Literature Map
  • Similar Papers
Abstract
Translate article icon Translate Article Star icon

The usability of stacking-based ensemble learning model in crime prediction: a systematic review

Similar Papers
  • Research Article
  • 10.1155/acis/5211419
Predicting Residential Energy Consumption in South Africa Using Ensemble Models
  • Jan 1, 2025
  • Applied Computational Intelligence and Soft Computing
  • David Attipoe + 3 more

This study presents ensemble machine learning (ML) models for predicting residential energy consumption in South Africa. By combining the best features of individual ML models, ensemble models reduce the drawbacks of each model and improve prediction accuracy. We present four ensemble models: ensemble by averaging (EA), ensemble by stacking each estimator (ESE), ensemble by boosting (EB), and ensemble by voting estimator (EVE). These models are built on top of Random Forest (RF) and Decision Tree (DT). These base predictor models leverage historical energy consumption patterns to capture temporal intricacies, including seasonal variations and rolling averages. In addition, we employed feature engineering methodologies to further enhance their predictive abilities. The accuracy of each ensemble model was evaluated by assessing various performance indicators, including the mean squared error (MSE), mean absolute error (MAE), mean absolute percentage error (MAPE), and coefficient of determination R2. Overall, the findings illustrate the efficiency of ensemble learning models in providing accurate predictions for residential energy consumption. This study provides valuable insights for researchers and practitioners in predicting energy consumption in residential buildings and the benefits of using ensemble learning models in the building and energy research domains.

  • Abstract
  • Cite Count Icon 3
  • 10.1093/ehjdh/ztac076.2784
Cardiovascular disease risk prediction via machine learning using mental health data
  • Dec 22, 2022
  • European Heart Journal. Digital Health
  • M Dorraki + 8 more

BackgroundRobust and accurate risk prediction models are much needed in cardiovascular disease. It is well-known that mental health is associated with the risk of developing cardiovascular disease. It is unknown whether mental health markers can enhance existing risk prediction models for cardiovascular disease.PurposeThe main purpose of this study was to assess capability of mental health factors along with traditional risk factors to be used in cardiovascular predictive machine learning models, and to develop a combined machine learning approach using both traditional risk and psychological factors in 375,145 participants of the UK Biobank.MethodsA comprehensive Pearson correlation analysis is carried out on UK Biobank data. Subsequently, an ensemble model containing decision tree, random forest, XGBoost, support vector machine (SVM), and deep neural network (DNN) classification approaches was built to predict cardiovascular diseases (CVD) in UK Biobank participants. The model was first trained using traditional cardiovascular risk factors, and subsequently trained using a combination of cardiovascular risk and psychological factors.ResultsThe correlation analysis revealed that there is a correlation between CVD and mental health factors suggesting the potential of mental health application for machine learning models. Our ensemble machine learning model was able to predict CVD with an accuracy of 73.49% using CVD risk factors alone. However, by combining psychological factors with CVD risk factors in the training data, an improved accuracy of 95.70% was achieved. The accuracy and robustness of ensemble machine learning model outperformed any of five constituent learning algorithms alone.ConclusionsOur results suggest that mental health assessment data along with traditional risk factors provides a powerful, safe and affordable machine learning model enrichment that can be used for state-of-the-art prediction of CVD.Funding AcknowledgementType of funding sources: None.Figure 1. Overview of CVD + Mental risk modelFigure 2. Results from the ensemble model

  • Research Article
  • Cite Count Icon 33
  • 10.1007/s11270-020-04693-w
Performance Evaluation of Homogeneous and Heterogeneous Ensemble Models for Groundwater Salinity Predictions: a Regional-Scale Comparison Study
  • Jun 1, 2020
  • Water, Air, & Soil Pollution
  • Alvin Lal + 1 more

Accurate prediction of salinity concentration in the aquifer in response to fluctuating groundwater pumping pattern is an essential component of any coastal groundwater planning and management framework. Data-driven prediction models have been proved efficient in predicting groundwater salinity levels in coastal aquifers. The use of ensemble prediction models is known to be more accurate with robust prediction capabilities when compared with standalone prediction models. This study compares the performances of homogeneous and heterogeneous ensemble models for groundwater salinity predictions. A homogeneous ensemble model is composed of several standalone models of the same type (i.e. employs one machine learning tool) whereas a heterogeneous ensemble model is composed of several standalone models of different types (i.e. employs multiple machine learning tools). Specifically, homogeneous and heterogeneous ensemble models of various standalone machine learning tools such as artificial neural network (ANN), genetic programming (GP), support vector regression (SVR), and Gaussian process regression (GPR) are developed to predict groundwater salinity concentrations in a small Pacific island coastal aquifer system. Standalone and ensemble prediction models are trained and validated using identical pumping and resulting salinity concentration datasets obtained by solving numerical 3D transient density-dependent coastal aquifer flow and transport model. After validation, the ensemble models are used to predict salinity concentration at selected monitoring wells in the modelled aquifer under variable groundwater pumping conditions. Prediction capabilities of the developed ensemble models are quantified using standard statistical procedures. The performance evaluation result suggested that the predictive capabilities of the developed standalone prediction models (ANN, GP, SVR, and GPR) were comparable with the numerical groundwater variable density-dependent flow and salt transport model. However, GPR standalone models had better prediction capabilities when compared with the other standalone models. Also, SVR and GPR standalone models were more efficient (i.e. took less computational training time) than other standalone models. In terms of ensemble models, the performance of the homogeneous GPR ensemble model was established to be superior to other homogeneous and heterogeneous ensemble models. The homogeneous GPR ensemble model was favoured both in terms of efficiency. Overall, based on the limited performance evaluation result, GPR homogeneous model was considered to be the best prediction model when compared with all the standalone models, other homogeneous ensemble model, and the heterogeneous ensemble model. Therefore, it can be utilised as a reliable groundwater salinity prediction tool and also used as an approximate simulator in coupled simulation-optimization models needed for prescribing optimal groundwater management strategies.

  • PDF Download Icon
  • Research Article
  • Cite Count Icon 16
  • 10.3389/frwa.2023.1305998
Investigating machine learning and ensemble learning models in groundwater potential mapping in arid region: case study from Tan-Tan water-scarce region, Morocco
  • Dec 13, 2023
  • Frontiers in Water
  • Abdessamad Jari + 7 more

Groundwater resource management in arid regions has a critical importance for sustaining human activities and ecological systems. Accurate mapping of groundwater potential plays a vital role in effective water resource planning. This study investigates the effectiveness of machine learning models, including Random Forest (RF), Adaboost, K-Nearest Neighbors (KNN), and Gaussian Process in groundwater potential mapping (GWPM) in the Tan-Tan arid region, Morocco. Fourteen groundwater conditional factors were considered following multicollinearity test, including topographical, hydrological, climatic, and geological factors. Additionally, point data with 174 sites indicative of groundwater occurrences were incorporated. The groundwater inventory data underwent random partitioning into training and testing datasets at three different ratios: 55/45%, 65/35%, and 75/25%. Ultimately, a comprehensive ranking of the 13 models, encompassing both individual and ensemble models, was determined using the prioritization rank technique. The results revealed that ensemble learning (EL) models, particularly RF and Adaboost (RF-Adaboost), outperformed individual models in groundwater potential mapping. Based on accuracy assessment using the validation dataset, the RF-Adaboost EL results yielded an Area Under the Receiver Operating characteristic Curve (AUROC) and Overall Accuracy (OA) of 94.02 and 94%, respectively. Ensemble models have been effectively applied to integrate 14 factors, capturing their intricate interrelationships, and thereby enhancing the accuracy and robustness of groundwater prediction in the Tan-Tan water-scarce region. Among the natural factors, the current study identified lithology, structural elements (such as faults and tectonic lineaments), and land use as significant contributors to groundwater potential. However, the critical characteristics of the study area showing a coastal position as well as a low background in groundwater prospectivity (low borehole points) are challenging in GWPM. The findings highlight the importance of the significant factors in assessing and managing groundwater resources in arid regions. Moreover, this study makes a contribution to the management of groundwater resources by demonstrating the effectiveness of ensemble learning algorithms in the groundwater potential mapping (GWPM) in arid regions.

  • Research Article
  • Cite Count Icon 5
  • 10.63125/edxgjg56
FORECASTING FUTURE INVESTMENT VALUE WITH MACHINE LEARNING, NEURAL NETWORKS, AND ENSEMBLE LEARNING: A META-ANALYTIC STUDY
  • Jun 1, 2022
  • Review of Applied Science and Technology
  • Kutub Uddin Apu + 3 more

This meta-analytic study investigates the effectiveness of machine learning (ML), neural networks (NN), and ensemble learning models in forecasting future investment value across diverse financial markets. Using PRISMA 2020 guidelines, 108 peer-reviewed articles published between 2012 and 2022 were systematically selected from databases including Scopus, Web of Science, and IEEE Xplore. The study synthesizes empirical findings on model performance, feature engineering, and algorithmic robustness to evaluate predictive accuracy, generalizability, and practical applicability. Results indicate that neural networks—particularly deep learning architectures such as LSTM and CNN—demonstrate superior performance in capturing nonlinear patterns and temporal dependencies in financial time series data. Ensemble models such as Random Forest, XGBoost, and hybrid frameworks (e.g., stacking, bagging, boosting) consistently outperform standalone ML models in terms of accuracy, stability, and resistance to overfitting. Approximately 34% of reviewed studies integrated macroeconomic indicators, technical indicators, and sentiment analysis to enhance feature richness, while 28% adopted multi-asset forecasting involving equities, cryptocurrencies, and derivatives. Performance metrics such as RMSE, MAPE, and R² revealed that ensemble and deep learning models achieve up to 20–30% improvement in predictive reliability compared to traditional statistical models like ARIMA and linear regression. The review also highlights a growing emphasis on model interpretability, with techniques like SHAP and LIME being applied in 18% of studies to support explainability in high-stakes investment decisions. However, challenges remain in model transparency, computational complexity, and adaptability across volatile market conditions. Compared to earlier literature, this study reflects a paradigm shift from linear forecasting models to adaptive, data-driven approaches supported by AI technologies. The findings underscore the transformative potential of ML, NNs, and ensemble models in investment forecasting while calling for continued research into scalable, explainable, and risk-aware deployment strategies for real-world financial environments.

  • Research Article
  • Cite Count Icon 159
  • 10.1016/j.geodrs.2020.e00256
Digital mapping of soil organic carbon using ensemble learning model in Mollisols of Hyrcanian forests, northern Iran
  • Feb 5, 2020
  • Geoderma Regional
  • Samaneh Tajik + 2 more

Digital mapping of soil organic carbon using ensemble learning model in Mollisols of Hyrcanian forests, northern Iran

  • Research Article
  • 10.38094/jastt62264
Feature-Based Child Mortality Prediction Using Ensemble and Traditional Machine Learning Models
  • Aug 8, 2025
  • Journal of Applied Science and Technology Trends
  • Aruna Sampathirao + 2 more

Child mortality is a big problem around the world, especially in low- and middle-income nations where there are big differences in health care and social conditions. This investigation seeks to create a predictive model for child mortality and pinpoint the key factors that significantly contribute to it, employing machine learning (ML) methodologies. The dataset includes various features such as parental age, maternal education, birth weight, wealth index, and access to healthcare services. Thirteen machine learning classifiers were used, categorized into four model groups: Traditional Models (Logistic Regression, K-Nearest Neighbors, Support Vector Machine, Naive Bayes), Tree-Based Models (Decision Tree, Random Forest, Extra Trees), Boosting Models (AdaBoost, Gradient Boosting, XGBoost), and Ensemble Learning Models (Soft Voting, Hard Voting, Stacking). The efficacy of each model was assessed using classification metrics, including Accuracy, Precision, Recall, and F1-Score within a 10-fold cross-validation framework to guarantee robustness. Results indicate that ensemble models, particularly AdaBoost, achieved the highest predictive accuracy, with perfect scores across all metrics (1.00). XGBoost and Stacking also demonstrated strong and consistent performance. The findings indicate that ensemble learning methods are effective in predicting child mortality and can assist policymakers and healthcare planners in identifying high-risk populations and implementing targeted interventions to reduce child mortality.

  • Research Article
  • Cite Count Icon 1
  • 10.1108/jfc-08-2024-0264
Detecting financial statement fraud using new ensemble learning: evidence during the COVID-19 pandemic in Indonesia
  • Mar 25, 2025
  • Journal of Financial Crime
  • Moh Riskiyadi

Purpose This study aims to propose a new ensemble learning model and compare its performance with other ensemble models to obtain the best model for detecting financial statement fraud during the COVID-19 pandemic. Design/methodology/approach This study uses a quantitative approach, using secondary data from financial reports, annual reports, regulatory reports and other information on the internet. It focuses on all companies listed on the Indonesia Stock Exchange from 2020 to 2023. The independent variables in this study use financial and nonfinancial variables. In contrast, the target variable for fraudulent financial reports is based on sanctions from regulators and the company’s special supervisory status. Findings This study results show that the ensemble blending model performs best in detecting financial statement fraud compared to the ensemble model that construct it. Research limitations/implications This study sets ensemble learning to default settings. Setting certain conditions can further improve the performance of ensemble learning models. Practical implications This study can broaden the insights of practitioners, academics, investors, regulators, stakeholders and corporate finance experts into detecting financial report fraud. Originality/value This study proposes a new ensemble learning model that previous studies have not discussed. This ensemble learning model performs best compared to other ensemble learning models.

  • Research Article
  • 10.3760/cma.j.cn112139-20240411-00180
Construction and verification of pancreatic fistula risk prediction model after pancreaticoduodenectomy based on ensemble machine learning
  • Oct 1, 2024
  • Zhonghua wai ke za zhi [Chinese journal of surgery]
  • S B Cheng + 8 more

Objective: To construct an ensemble machine learning model for predicting the occurrence of clinically relevant postoperative pancreatic fistula (CR-POPF) after pancreaticoduodenectomy and evaluate its application value. Methods: This is a research on predictive model. Clinical data of 421 patients undergoing pancreaticoduodenectomy in the Department of Pancreatic Surgery,Union Hospital, Tongji Medical College,Huazhong University of Science and Technology from June 2020 to May 2023 were retrospectively collected. There were 241 males (57.2%) and 180 females (42.8%) with an age of (59.7±11.0)years (range: 12 to 85 years).The research objects were divided into training set (315 cases) and test set (106 cases) by stratified random sampling in the ratio of 3∶1. Recursive feature elimination is used to screen features,nine machine learning algorithms are used to model,three groups of models with better fitting ability are selected,and the ensemble model was constructed by Stacking algorithm for model fusion. The model performance was evaluated by various indexes,and the interpretability of the optimal model was analyzed by Shapley Additive Explanations(SHAP) method. The patients in the test set were divided into different risk groups according to the prediction probability (P) of the alternative pancreatic fistula risk score system (a-FRS). The a-FRS score was validated and the predictive efficacy of the model was compared. Results: Among 421 patients,CR-POPF occurred in 84 cases (20.0%). In the test set,the Stacking ensemble model performs best,with the area under the curve (AUC) of the subject's work characteristic curve being 0.823,the accuracy being 0.83,the F1 score being 0.63,and the Brier score being 0.097. SHAP summary map showed that the top 9 factors affecting CR-POPF after pancreaticoduodenectomy were pancreatic duct diameter,CT value ratio,postoperative serum amylase,IL-6,body mass index,operative time,albumin difference before and after surgery,procalcitonin and IL-10. The effects of each feature on the occurrence of CR-POPF after pancreaticoduodenectomy showed a complex nonlinear relationship. The risk of CR-POPF increased when pancreatic duct diameter<3.5 mm,CT value ratio<0.95,postoperative serum amylase concentration>150 U/L,IL-6 level>280 ng/L,operative time>350 minutes,and albumin decreased by more than 10 g/L. The AUC of a-FRS in the test set was 0.668,and the prediction performance of a-FRS was lower than that of the Stacking ensemble machine learning model. Conclusion: The ensemble machine learning model constructed in this study can predict the occurrence of CR-POPF after pancreaticoduodenectomy,and has the potential to be a tool for personalized diagnosis and treatment after pancreaticoduodenectomy.

  • Research Article
  • Cite Count Icon 1
  • 10.1093/eurheartj/ehac544.2784
Cardiovascular disease risk prediction via machine learning using mental health data
  • Oct 3, 2022
  • European Heart Journal
  • M Dorraki + 8 more

Cardiovascular disease risk prediction via machine learning using mental health data

  • Research Article
  • Cite Count Icon 27
  • 10.1016/j.buildenv.2022.109462
Comparative analysis of thermal preference prediction performance in different conditions using ensemble learning models based on ASHRAE Comfort Database II
  • Sep 1, 2022
  • Building and Environment
  • Yan Bai + 2 more

Comparative analysis of thermal preference prediction performance in different conditions using ensemble learning models based on ASHRAE Comfort Database II

  • Research Article
  • Cite Count Icon 2
  • 10.1016/j.jdent.2024.105467
Development of a survey-based stacked ensemble predictive model for autonomy preferences in patients with periodontal disease
  • Nov 18, 2024
  • Journal of Dentistry
  • So-Hae Oh + 5 more

Development of a survey-based stacked ensemble predictive model for autonomy preferences in patients with periodontal disease

  • Research Article
  • Cite Count Icon 33
  • 10.1016/j.fuel.2021.121975
An ensemble deep learning model for exhaust emissions prediction of heavy oil-fired boiler combustion
  • Sep 15, 2021
  • Fuel
  • Zhezhe Han + 5 more

An ensemble deep learning model for exhaust emissions prediction of heavy oil-fired boiler combustion

  • Research Article
  • 10.1108/ohi-02-2025-0063
The role of urban environment maintenance in crime prevention: a machine learning analysis of street crime in Manhattan
  • Mar 27, 2026
  • Open House International
  • Woocheol Kim + 2 more

Purpose Urban environment maintenance is a critical component of Crime Prevention Through Environmental Design (CPTED), yet its direct impact on crime reduction remains underexplored. Grounded in the Broken Windows Theory, this study investigates the relationship between urban maintenance factors and street crime in Manhattan, New York City. Design/methodology/approach This research employs a machine learning approach to analyze crime occurrences in relation to urban maintenance indicators. Using a grid-based spatial analysis (250 × 250 meters), we assess over 18,000 citizen-reported complaints from the 311 system spanning five years. Ensemble learning models, including Random Forest and XGBoost, were applied to predict crime occurrence, with SHAP (SHapley Additive exPlanations) analysis used to interpret feature importance. Findings The results reveal that environmental maintenance factors influencing crime vary significantly depending on local urban conditions. In high-crime areas, specific factors such as illegal parking, homeless presence, and abandoned vehicles exhibit stronger associations with crime rates, whereas noise and sanitation concerns play a more prominent role in low-crime areas. These findings suggest that crime prevention strategies should incorporate region-specific maintenance interventions rather than a one-size-fits-all approach. Originality/value This study advances the understanding of urban crime prevention by integrating machine learning-based predictive modeling with spatial environmental analysis. By identifying key maintenance factors contributing to crime variations at a granular level, the research provides actionable insights for policymakers, urban planners, and law enforcement agencies seeking to enhance urban safety through targeted environmental management.

  • PDF Download Icon
  • Research Article
  • Cite Count Icon 68
  • 10.3389/fpubh.2022.892371
Artificial Intelligence-Based Ensemble Learning Model for Prediction of Hepatitis C Disease.
  • Apr 27, 2022
  • Frontiers in Public Health
  • Michael Onyema Edeh + 6 more

Machine learning algorithms are excellent techniques to develop prediction models to enhance response and efficiency in the health sector. It is the greatest approach to avoid the spread of hepatitis C, especially injecting drugs, is to avoid these behaviors. Treatments for hepatitis C can cure most patients within 8 to 12 weeks, so being tested is critical. After examining multiple types of machine learning approaches to construct the classification models, we built an AI-based ensemble model for predicting Hepatitis C disease in patients with the capacity to predict advanced fibrosis by integrating clinical data and blood biomarkers. The dataset included a variety of factors related to Hepatitis C disease. The training data set was subjected to three machine-learning approaches and the validated data was then used to evaluate the ensemble learning-based prediction model. The results demonstrated that the proposed ensemble learning model has been observed ad more accurate compared to the existing Machine learning algorithms. The Multi-layer perceptron (MLP) technique was the most precise learning approach (94.1% accuracy). The Bayesian network was the second-most accurate learning algorithm (94.47% accuracy). The accuracy improved to the level of 95.59%. Hepatitis C has a significant frequency globally, and the disease's development can result in irreparable damage to the liver, as well as death. As a result, utilizing AI-based ensemble learning model for its prediction is advantageous in curbing the risks and improving treatment outcome. The study demonstrated that the use of ensemble model presents more precision or accuracy in predicting Hepatitis C disease instead of using individual algorithms. It also shows how an AI-based ensemble model could be used to diagnose Hepatitis C disease with greater accuracy.

Save Icon
Up Arrow
Open/Close
Notes

Save Important notes in documents

Highlight text to save as a note, or write notes directly

You can also access these Documents in Paperpal, our AI writing tool

Powered by our AI Writing Assistant