Accelerate Literature Icon
Want to do a literature review? Try our new Literature Review workflow

Design of a Predictive Model Using AI and GIS for the Management and Optimization of Resources in Forest Fires in Eastern Michoacán

  • Abstract
  • Literature Map
  • Similar Papers
Abstract
Translate article icon Translate Article Star icon

This study describes a design and validation proposal for an artificial intelligence–based predictive model to estimate the area affected by forest fires and support the optimisation of operational resource use in the eastern region of Michoacán, Mexico. Using a dataset comprising 930 historical records (2015–2024), three supervised learning algorithms were evaluated: Multilayer Perceptron (MLP), Random Forest (RF), and XGBoost. The MLP model, optimised using L2 regularisation, Dropout, and cross-validation, achieved the highest performance, with a coefficient of determination (R²) of 0.76, compared with RF (0.57), XGBoost (0.51), and a weighted average ensemble model (0.60). The results were integrated into interactive cartographic platforms using QGIS and Leaflet JS, enabling the visualisation of areas with higher fire susceptibility and the prioritisation of interventions. A replicable tool with low computational requirements is put forward as suitable for institutional contexts with limited resources. Smart citations: https://scite.ai/reports/10.61467/2007.1558.2026.v17i1.1181Dimensions.Open Alex.

Similar Papers
  • Research Article
  • Cite Count Icon 11
  • 10.3808/jeil.201900010
Short-Term Wastewater Influent Prediction Based on Random Forests and Multi-Layer Perceptron
  • Jan 1, 2019
  • Journal of Environmental Informatics Letters
  • P Zhou + 4 more

Influent flow rate is a crucial parameter closely related to the plant-wide control of wastewater treatment plants (WWTPs). In this study, a random forest (RF) model and a multi-layer perceptron (MLP) model are developed for hourly influent flow rate prediction at a confidential WWTP in Canada. Both models perform well on predicting influent flow rate one-step ahead. The coefficient of determination (R2) values of MLP and RF for the testing data set are 0.927 and 0.925, respectively. Furthermore, the multi-step ahead prediction accuracy of the proposed models is discussed. To improve the multi-step ahead prediction accuracy of the RF model, time-tag information is transformed to numerical values and then fed into the RF model as input. The R2 values of the RF model for the testing data set with and without time-tag information are 0.334 and 0.811, respectively. The results show that the RF model's performance for multi- step ahead prediction is heavily affected by the time-tag information. Including time-tag information as input could dramatically improve the multi-step ahead prediction accuracy. In this study, the RF model shows more robust performance than the MLP model on solving short-term wastewater influent prediction problems. Influent flow rate is a crucial parameter closely related to the plant-wide control of wastewater treatment plants (WWTPs). In this study, a random forest (RF) model and a multi-layer perceptron (MLP) model are developed for hourly influent flow rate prediction at a confidential WWTP in Canada. Both models perform well on predicting influent flow rate one-step ahead. The coefficient of determination (R2) values of MLP and RF for the testing data set are 0.927 and 0.925, respectively. Furthermore, the multi-step ahead prediction accuracy of the proposed models is discussed. To improve the multi-step ahead prediction accuracy of the RF model, time-tag information is transformed to numerical values and then fed into the RF model as input. The R2 values of the RF model for the testing data set with and without time-tag information are 0.334 and 0.811, respectively. The results show that the RF model's performance for multi-step ahead prediction is heavily affected by the time-tag information. Including time-tag information as input could dramatically improve the multi-step ahead prediction accuracy. In this study, the RF model shows more robust performance than the MLP model on solving short-term wastewater influent prediction problems.

  • Research Article
  • 10.1080/10549811.2026.2660113
Analysis of Forest Fires in Türkiye Using Machine Learning and Artificial Neural Networks
  • Apr 18, 2026
  • Journal of Sustainable Forestry
  • Alaattin Parlakkılıç

This study analyzes forest fires in Turkey between 2012 and 2022 using machine learning and artificial neural network methods, considering their timing, causes, burned area size, and spatial distribution. Official and publicly available data from the Turkish Orman Genel Müdürlüğü were used as the dataset. Decision Tree, Random Forest, Support Vector Machines, Multilayer Perceptron, and Long Short-Term Memory models were applied and comparatively evaluated. The performance of the classification models was evaluated using accuracy and F1 score, while temporal predictions were assessed using the Root Mean Square Error metric. According to the results, the Multilayer Perceptron model showed the highest success in predicting fire causes and burned area size classes (accuracy = 86%, F1 = 0.83). The Random Forest model similarly demonstrated strong performance with 85% accuracy and an F1 score of 0.82. The Long Short-Term Memory model, used for temporal predictions, achieved the lowest error value (Root Mean Square Error = 0.12) in estimating the annual and monthly number of fires and the amount of burned area, successfully capturing the variability, especially in years with excessive fires. Overall, the findings show that the Long Short-Term Memory model is superior in modeling temporal dependencies, while the Random Forest and Multilayer Perceptron models provide reliable results in classification problems. In conclusion, machine learning and artificial neural networks can be effectively used in the development of early warning systems, risk mapping, and decision support mechanisms for forest fires.

  • Research Article
  • Cite Count Icon 7
  • 10.2147/clep.s505966
Explainable Prediction of Long-Term Glycated Hemoglobin Response Change in Finnish Patients with Type 2 Diabetes Following Drug Initiation Using Evidence-Based Machine Learning Approaches.
  • Mar 1, 2025
  • Clinical epidemiology
  • Gunjan Chandra + 7 more

This study applied machine learning (ML) and explainable artificial intelligence (XAI) to predict changes in HbA1c levels, a critical biomarker for monitoring glycemic control, within 12 months of initiating a new antidiabetic drug in patients diagnosed with type 2 diabetes. It also aimed to identify the predictors associated with these changes. Electronic health records (EHR) from 10,139 type 2 diabetes patients in North Karelia, Finland, were used to train models integrating randomized controlled trial (RCT)-derived HbA1c change values as predictors, creating offset models that integrate RCT insights with real-world data. Various ML models-including linear regression (LR), multi-layer perceptron (MLP), ridge regression (RR), random forest (RF), and XGBoost (XGB)-were evaluated using R² and RMSE metrics. Baseline models used data at or before drug initiation, while follow-up models included the first post-drug HbA1c measurement, improving performance by incorporating dynamic patient data. Model performance was also compared to expected HbA1c changes from clinical trials. Results showed that ML models outperform RCT model, while LR, MLP, and RR models had comparable performance, RF and XGB models exhibited overfitting. The follow-up MLP model outperformed the baseline MLP model, with higher R² scores (0.74, 0.65) and lower RMSE values (6.94, 7.62), compared to the baseline model (R²: 0.52, 0.54; RMSE: 9.27, 9.50). Key predictors of HbA1c change included baseline and post-drug initiation HbA1c values, fasting plasma glucose, and HDL cholesterol. Using EHR and ML models allows for the development of more realistic and individualized predictions of HbA1c changes, accounting for more diverse patient populations and their heterogeneous nature, offering more tailored and effective treatment strategies for managing T2D. The use of XAI provided insights into the influence of specific predictors, enhancing model interpretability and clinical relevance. Future research will explore treatment selection models.

  • PDF Download Icon
  • Research Article
  • Cite Count Icon 28
  • 10.3390/agronomy11030532
Machine-Learning Approach Using SAR Data for the Classification of Oil Palm Trees That Are Non-Infected and Infected with the Basal Stem Rot Disease
  • Mar 12, 2021
  • Agronomy
  • Izrahayu Che Hashim + 4 more

Basal stem rot disease (BSR) in oil palm plants is caused by the Ganoderma boninense (G. boninense) fungus. BSR is a major disease that affects oil palm plantations in Malaysia and Indonesia. As of now, the only available sustaining measure is to prolong the life of oil palm trees since there has been no effective treatment for the BSR disease. This project used an ALOS PALSAR-2 image with dual polarization, Horizontal transmit and Horizontal receive (HH) and Horizontal transmit and Vertical receive (HV). The aims of this study were to (1) identify the potential backscatter variables; and (2) examine the performance of machine learning (ML) classifiers (Multilayer Perceptron (MLP) and Random Forest (RF) to classify oil palm trees that are non-infected and infected by G. boninense. The sample size consisted of 55 uninfected trees and 37 infected trees. We used the imbalance data approach (Synthetic Minority Over-Sampling Technique (SMOTE) in these classifications due to the differing sample sizes. The result showed backscatter variable HV had a higher correct classification for the G. boninense non-infected and infected oil palm trees for both classifiers; the MLP classifier model had a robust success rate, which correctly classified 100% for non-infected and 91.30% for infected G. boninense, and RF had a robust success rate, which correctly classified 94.11% for non-infected and 91.30% for infected G. boninense. In terms of model performance using the most significant variables, HV, the MLP model had a balanced accuracy (BCR) of 95.65% compared to 92.70% for the RF model. Comparison between the MLP model and RF model for the receiver operating characteristics (ROC) curve region, (AUC) gave a value of 0.92 and 0.95, respectively, for the MLP and RF models. Therefore, it can be concluded by using only the HV polarization, that both the MLP and RF can be used to predict BSR disease with a relatively high accuracy.

  • PDF Download Icon
  • Research Article
  • Cite Count Icon 37
  • 10.1080/24749508.2019.1610841
Model-based soil temperature estimation using climatic parameters: the case of Azerbaijan Province, Iran
  • Apr 25, 2019
  • Geology, Ecology, and Landscapes
  • Parveen Sihag + 3 more

ABSTRACTEstimating soil temperature (ST) profile is identified as essential knowledge for plants, crop growth, and germination in all agriculture regions. In this study, daily soil temperature (DST) was modeled using Multilayer perceptron (MLP) model, Gaussian Process (GP), Random Forest (RF), and the M5P model methods for estimating and comparing DST in arid regions. The data selected to test the proposed models are obtained from two stations in Tabriz and Ahar, located in the Azerbaijan province of Iran. Input dataset includes air temperature, relative humidity, wind speed, and sunshine as dependent parameters, whereas ST at depths of 5 cm was selected for the target in model development. The results show the MLP works better than GP-, RF-, and M5P-based models in estimating the DST, with excellent performance indicators such as the mean absolute error, root mean square error, and coefficient of correlation. Results showed that the MLP model with RMSE = 3.2626°C was more suitable than other models in ST estimation 2 days ahead for Tabriz station. Also, in Ahar, MLP with RMSE = 6.3332°C was more suitable than GP-, RF-, and M5P-based models for estimating DST. As a conclusion, the developed MLP is recommended for estimating the DST profiles.

  • Research Article
  • Cite Count Icon 39
  • 10.1016/j.enconman.2024.118808
Machine learning development to predict the electrical efficiency of photovoltaic-thermal (PVT) collector systems
  • Jul 23, 2024
  • Energy Conversion and Management
  • Hossein Gharaee + 2 more

According to the increasing rate of renewable energy use, especially solar energy, photovoltaic-thermal (PVT) solar panels are getting more attention and they have become significantly important for extracting heat and electricity from solar energy. In this regard, the theoretical equations used to calculate the PVT performance lack accuracy, and many uncertainties and calculation impediments result in deviation from the experimental outputs. Therefore, machine learning (ML) methods can overcome these disadvantages and can be used to predict PVT efficiency more reliably and effectively. In this study, three ML methods, namely, multilayer perceptron (MLP), random forest (RF), and support vector regression (SVR) were implemented to build trained models predicting the electrical efficiency of PVTs. These models are based on more than 380 datasets that were extracted from the literature where all PVT benches use water as their working fluid. The models considered mass flow rate, solar radiation, ambient temperature, wind speed, fluid inlet temperature, PVT surface area, and pipe inner diameter as input variables. The optimized models were attained through rigorous hyperparameters variations through random, coarse and fine grid searches. The results showed that the RF model has a root mean squared error (RMSE) of 0.2233, 0.4638, and 0.671 for training, testing, and validation datasets respectively, and training coefficient of determination (R-square), with a value of 0.9862, providing relatively high prediction accuracy for the electrical efficiency. Next, the MLP model with a testing R-square value of 0.8775 and the SVR model with a testing R-square value of 0.7639 can predict electrical efficiencies. Further, the output results are visualized using explainable artificial intelligence (AI) method such as SHapley Additive exPlanations (SHAP) which declared that mass flow rate, pipe inner diameter, and wind speed are the most effective variables in the RF model. For validation purposes, 46 datasets characterized with at least one out-domain parameter are used to test the performance of our models. For out-of-range input variables, the RF model most accurately predicted the electrical efficiency for cell absorptance and inlet fluid temperature variables, outperforming the MLP and SVR models. Meanwhile, the SVR model showed superior accuracy over both RF and MLP when dealing with datasets featuring solar radiation values beyond the training domain.

  • PDF Download Icon
  • Research Article
  • Cite Count Icon 103
  • 10.3390/w14223743
Application of Machine Learning Techniques for the Estimation of the Safety Factor in Slope Stability Analysis
  • Nov 18, 2022
  • Water
  • Yaser Ahangari Nanehkaran + 6 more

Slope stability is the most important stage in the stabilization process for different scale slopes, and it is dictated by the factor of safety (FS). The FS is a relationship between the geotechnical characteristics and the slope behavior under various loading conditions. Thus, the application of an accurate procedure to estimate the FS can lead to a fast and precise decision during the stabilization process. In this regard, using computational models that can be operated accurately is strongly needed. The performance of five different machine learning models to predict the slope safety factors was investigated in this study, which included multilayer perceptron (MLP), support vector machines (SVM), k-nearest neighbors (k-NN), decision tree (DT), and random forest (RF). The main objective of this article is to evaluate and optimize the various machine learning-based predictive models regarding FS calculations, which play a key role in conducting appropriate stabilization methods and stabilizing the slopes. As input to the predictive models, geo-engineering index parameters, such as slope height (H), total slope angle (β), dry density (γd), cohesion (c), and internal friction angle (φ), which were estimated for 70 slopes in the South Pars region (southwest of Iran), were considered to predict the FS properly. To prepare the training and testing data sets from the main database, the primary set was randomly divided and applied to all predictive models. The predicted FS results were obtained for testing (30% of the primary data set) and training (70% of the primary data set) for all MLP, SVM, k-NN, DT, and RF models. The models were verified by using a confusion matrix and errors table to conclude the accuracy evaluation indexes (i.e., accuracy, precision, recall, and f1-score), mean squared error (MSE), mean absolute error (MAE), and root mean square error (RMSE). According to the results of this study, the MLP model had the highest evaluation with a precision of 0.938 and an accuracy of 0.90. In addition, the estimated error rate for the MLP model was MAE = 0.103367, MSE = 0.102566, and RMSE = 0.098470.

  • PDF Download Icon
  • Research Article
  • Cite Count Icon 19
  • 10.3390/w7062707
Spatial Disaggregation of Areal Rainfall Using Two Different Artificial Neural Networks Models
  • Jun 5, 2015
  • Water
  • Sungwon Kim + 1 more

The objective of this study is to develop artificial neural network (ANN) models, including multilayer perceptron (MLP) and Kohonen self-organizing feature map (KSOFM), for spatial disaggregation of areal rainfall in the Wi-stream catchment, an International Hydrological Program (IHP) representative catchment, in South Korea. A three-layer MLP model, using three training algorithms, was used to estimate areal rainfall. The Levenberg–Marquardt training algorithm was found to be more sensitive to the number of hidden nodes than were the conjugate gradient and quickprop training algorithms using the MLP model. Results showed that the networks structures of 11-5-1 (conjugate gradient and quickprop) and 11-3-1 (Levenberg-Marquardt) were the best for estimating areal rainfall using the MLP model. The networks structures of 1-5-11 (conjugate gradient and quickprop) and 1-3-11 (Levenberg–Marquardt), which are the inverse networks for estimating areal rainfall using the best MLP model, were identified for spatial disaggregation of areal rainfall using the MLP model. The KSOFM model was compared with the MLP model for spatial disaggregation of areal rainfall. The MLP and KSOFM models could disaggregate areal rainfall into individual point rainfall with spatial concepts.

  • Research Article
  • 10.1038/s41598-026-56583-6
Multimodal modeling framework and characteristic analysis of a new air-source heat pump heating system.
  • Jun 3, 2026
  • Scientific reports
  • Mingzhi Jiang + 5 more

Heating systems that combine air-source heat pumps with phase-change energy storage tanks can effectively utilize off-peak electricity and enhance energy storage efficiency. However, in actual industrial building heating operations, the system's thermal characteristics are influenced by multiple factors, making it challenging for traditional single-factor prediction models to accurately capture their dynamic behavior. To address this issue, this paper develops a novel prediction model called XMS (XGBoost, MLP, and Stacking model). First, heating data from an operational air-source heat pump system coupled with a phase-change energy storage tank are collected. Second, the XMS model is proposed, based on a stacking ensemble strategy. It integrates XGBoost and a multi-layer perceptron (MLP) as base learners and employs a meta-learner to perform secondary modeling of their outputs. Finally, the XMS model is compared with a traditional MLP model. The results indicate that the XMS model demonstrates significantly superior predictive performance compared to the MLP model, achieving a 22.5% reduction in root mean square error (RMSE) and a 10.1% increase in the coefficient of determination (R²). The XMS model better captures the nonlinear and multivariate coupling characteristics of the heating system. This study provides an efficient and reliable modeling method for load forecasting and control optimization in air-source heat pump phase-change energy storage heating systems, offering valuable guidance for intelligent heating control technology.

  • Conference Article
  • 10.2118/229716-ms
Innovative AI Agent for Real-Time Drill Bit Selection Optimization
  • Nov 3, 2025
  • Matin Shahin + 7 more

Objectives/Scope Optimizing drill bit selection constitutes a critical challenge within the domain of oil and gas drilling, with substantial implications for operational efficiency and cost-effectiveness. Traditional methodologies, often reliant on empirical guidelines or experiential knowledge, are prone to biases and inadequacies in addressing the complexities of contemporary drilling data. This study presents an innovative artificial intelligence (AI) agent that leverages advanced machine learning techniques, specifically utilizing Multi-Layer Perceptron (MLP) and Random Forest (RF) algorithms, to enhance the real-time selection of drill bits as informed by the International Association of Drilling Contractors (IADC) Code. Methods, Procedures, Process The AI agent independently employs both MLP and RF models to analyze a comprehensive dataset comprising over 1 million drilling records sourced from Middle Eastern oilfields. Key input parameters include Measured Depth (MD), Weight-On-Bit (WOB), Rotational Speed (RPM), Rotary Torque (TQ), Rate of Penetration (ROP), Pump Pressure, Flow Rate, Mud Weight (MW), True Vertical Depth (TVD), Bit Size, Drilled Interval, Total Flow Area (TFA), Jet Number, and Formation Types. The analysis targets the prediction of the IADC Code. Results, Observations, Conclusions Field implementation of the AI agent yielded significant improvements in the accuracy and consistency of drill bit selection outcomes. The MLP model achieved an impressive overall accuracy of 99.3%, with an F1 score of 96.0%, whereas the RF model attained an accuracy of 95.1% and an F1 score of 93.4%. Both models demonstrated robust precision and recall rates; however, the MLP model exhibited superior performance in accurately classifying IADC Codes, particularly for less frequent categories. The probabilistic outputs produced by these models empower drilling engineers to make informed decisions based on quantifiable confidence metrics, thereby enhancing the decision-making process in dynamic drilling environments. Novel/Additive Information This study introduces a novel AI agent framework that systematically employs MLP and RF methodologies to provide real-time, autonomous, and context-sensitive recommendations for drill bit selection. This approach signifies a substantial advancement in digital drilling optimization, establishing a new benchmark for intelligent decision support systems within the petroleum industry and facilitating the ongoing digital transformation of drilling operations.

  • Research Article
  • Cite Count Icon 18
  • 10.1016/j.energy.2023.129433
Modeling of asphaltic sludge formation during acidizing process of oil well reservoir using machine learning methods
  • Oct 21, 2023
  • Energy
  • Sina Shakouri + 1 more

Modeling of asphaltic sludge formation during acidizing process of oil well reservoir using machine learning methods

  • Research Article
  • Cite Count Icon 3
  • 10.1002/ece3.71499
Predicting Aboveground Carbon Storage in Different Types of Forests in South Subtropical Regions Using Machine Learning Models.
  • May 1, 2025
  • Ecology and evolution
  • Jiarun Liu + 5 more

Motivated by the need to enhance the accuracy of forest aboveground carbon storage (ACS) assessments, this study aimed to explore the effectiveness of different machine learning models in predicting ACS across various subtropical forest types in southern China. The study was conducted in southern China, focusing on different types of subtropical forests. This region harbors several types of subtropical forests, which are rarely found at similar latitudes in the world. Variance inflation factor was employed to screen independent variables, resulting in the selection of 13 significant predictors. Four machine learning models-support vector machine (SVM), random forest (RF), multi-layer perceptron (MLP), and extreme gradient boosting (XGB)-were constructed to estimate carbon storage. Model performance was evaluated using root mean square error, coefficient of determination (R 2), and mean absolute error. The model with the best generalization ability was selected to calculate SHAP values for each predictor. The XGB model demonstrated superior performance across all forest types, with R 2 values ranging from 0.898 to 0.974. In mountainous evergreen broad-leaved forests, the prediction accuracy followed the order of XGB>MLP>SVM>RF. In valley rainforests, MLP showed the highest R 2 value, but with higher MAE and RMSE, making it the second-best choice. The RF model performed moderately, while the SVM model showed the poorest performance. The SHAP values indicated that maximum diameter at breast height, slope, mean DBH, species evenness, altitude, and maximum tree height had significant effects on ACS. XGB model exhibits the best prediction performance and strongest adaptability for estimating ACS in subtropical southern China forests. Additionally, the MLP model can serve as an effective model for assessing carbon storage in valley rainforests within this region. Machine learning methods provide valuable references for predicting and assessing ACS in different types of zonal forests.

  • Research Article
  • Cite Count Icon 7
  • 10.1080/10903127.2019.1610531
Risk Factors for Non-optimal Resource Utilization for Emergent Interfacility Transfers by Air Ambulance in Ontario
  • May 17, 2019
  • Prehospital Emergency Care
  • Brodie Nolan + 6 more

Background: The use of air ambulance to facilitate interfacility transfer has been associated with improved mortality; however, air ambulance is a limited resource and sometimes the optimal resource to transport a patient is unavailable. When a non-optimal resource is used there is an inherent delay and critically unwell patients may deteriorate as a result. This study aimed to identify risk factors associated with non-optimal resource utilization for adult patients undergoing emergent interfacility transport by air ambulance in Ontario, Canada. A secondary objective was to determine if non-optimal resource utilization was associated with deterioration in clinical status by measuring a delta rapid emergency medicine score (REMS). Methods: This was a retrospective cohort study of all emergent, adult interfacility transfers transported by air ambulance over a 5-year period in Ontario, Canada. Determination of optimal resource use was based on distances and historic time data for all sending-receiving facility pairs. A logistic regression model was used to explore patient, provider and institutional risk factors for non-optimal resource use. To explore the secondary objective a linear regression model was used to explore impact of non-optimal resource use on deltaREMS. Results: There were a total of 9,687 patients included in the study cohort, with 4,984 having an optimal resource use and 4,703 having non-optimal resource. The median delay in interfacility transfer caused by a non-optimal transfer strategy was 35.7 minutes. Patients who required mechanical ventilation (OR 1.13, p = 0.031) and or were transferred out of nursing stations had higher odds of non-optimal resource use (OR 2.84, p = 0.019). Paramedic level of care of advanced (OR 0.37, p = < 0.001) and critical care (OR 0.28, p = < 0.001) as well as spring season (OR 0.75, p = < 0.001) had lower odds of non-optimal resource utilization. Optimal resource utilization did not significantly affect delta REMS (beta coefficient 0.002, p = 0.64). Conclusions: Patients who required mechanical ventilation and were transferred out from a nursing station had higher odds of non-optimal resource utilization while patients that required advanced or critical care level of care and spring season had lower odds of non-optimal resource use. Additionally, non-optimal resource use for air ambulance interfacility transfers did not result in patient deterioration as measured by a delta REMS score.

  • Research Article
  • Cite Count Icon 50
  • 10.1016/j.artmed.2019.07.008
Optimizing neural networks for medical data sets: A case study on neonatal apnea prediction.
  • Jul 1, 2019
  • Artificial Intelligence in Medicine
  • Rudresh Deepak Shirwaikar + 5 more

Optimizing neural networks for medical data sets: A case study on neonatal apnea prediction.

  • Research Article
  • Cite Count Icon 39
  • 10.14710/ijred.2022.41451
Machine Learning Models Based on Random Forest Feature Selection and Bayesian Optimization for Predicting Daily Global Solar Radiation
  • Feb 1, 2022
  • International Journal of Renewable Energy Development
  • Mohamed Chaibi + 4 more

Prediction of daily global solar radiation with simple and highly accurate models would be beneficial for solar energy conversion systems. In this paper, we proposed a hybrid machine learning methodology integrating two feature selection methods and a Bayesian optimization algorithm to predict H in the city of Fez, Morocco. First, we identified the most significant predictors using two Random Forest methods of feature importance: Mean Decrease in Impurity (MDI) and Mean Decrease in Accuracy (MDA). Then, based on the feature selection results, ten models were developed and compared: (1) five standalone machine learning (ML) models including Classification and Regression Trees (CART), Random Forests (RF), Bagged Trees Regression (BTR), Support Vector Regression (SVR), and Multi-Layer Perceptron (MLP); and (2) the same models tuned by the Bayesian optimization (BO) algorithm: CART-BO, RF-BO, BTR-BO, SVR-BO, and MLP-BO. Both MDI and MDA techniques revealed that extraterrestrial solar radiation and sunshine duration fraction were the most influential features. The BO approach improved the predictive accuracy of MLP, CART, SVR, and BTR models and prevented the CART model from overfitting. The best improvements were obtained using the MLP model, where RMSE and MAE were reduced by 17.6% and 17.2%, respectively. Among the studied models, the SVR-BO algorithm provided the best trade-off between prediction accuracy (RMSE=0.4473kWh/m²/day, MAE=0.3381kWh/m²/day, and R²=0.9465), stability (with a 0.0033kWh/m²/day increase in RMSE), and computational cost.

Save Icon
Up Arrow
Open/Close
Notes

Save Important notes in documents

Highlight text to save as a note, or write notes directly

You can also access these Documents in Paperpal, our AI writing tool

Powered by our AI Writing Assistant