Univariate Monthly Rainfall Forecasting in Nigeria Using Multiple Statistical and Machine Learning Methods
This study evaluates statistical and machine learning models for monthly rainfall forecasting in Nigeria, finding SARIMA models, properly tuned, outperform ML approaches, especially in regions with minimal rainfall variability, achieving RMSE as low as 7.84mm and R2 up to 0.85, highlighting SARIMA's suitability for Nigeria's climate.
Abstract Precise long-term rainfall prediction is important for agricultural planning, climate resilience, and reducing disaster risk, particularly for countries like Nigeria with diverse regimes of rainfall. In this research, the potential of machine learning (ML) and statistical models to predict monthly univariate rainfall in 24 Nigerian stationswas evaluated. Model training employed historical rainfall data (1960–1999), while validation was carried out for 11 years (2000–2010). SARIMA ( p; d; q ) ( P; D; Q ) s models were used in Minitab ® , R, and Python, and the most important parameters ( p; d; q; P; D; Q ) were tuned manually and by using auto.arima(). ML models such as feedforward neural networks, adaptive neuro-fuzzy inference systems, support vector regression and random forest were utilized in MATLAB ® and R with hyperparameter-tuned models. Model performancewas evaluated in using statistics such as root mean square error ( RMSE ) and coefficient of determination ( r 2 ). SARIMA performed best in areas where rainfall variability was minimal. Nguru (12.03°N), the area with the lowest average monthly rainfall (35.71 mm), showed the highest SARIMA estimation with RMSE of as low as 7.84mm and r 2 of as high as 0.85. ML models underperformed in capturing seasonal dynamics. For instance, SVR failed to model temporal trends effectively, while random forest produced nearly constant outputs across all years. Adjustments to SARIMA parameters (e.g., setting seasonal differencing D = 0 or Q = 1) were essential in reducing unrealistic forecasts. The findings demonstrate that SARIMA, with proper tuning, is better suited for univariate rainfall forecasting in Nigeria than non-customized ML models. Forecast reliability strongly correlates with regional rainfall characteristics and model sensitivity to seasonality.
- Research Article
62
- 10.1016/j.jhydrol.2020.124759
- Feb 25, 2020
- Journal of Hydrology
On the complexities of sediment load modeling using integrative machine learning: Application of the great river of Loíza in Puerto Rico
- Research Article
- 10.1080/19942060.2026.2665857
- Dec 31, 2026
- Engineering Applications of Computational Fluid Mechanics
The Global Flood Awareness System (GloFAS) is a promising flood-modelling tool available for most river basins worldwide. However, it contains inherent biases, notably linked to its spatial resolution. This study proposes a multi-station ensemble modelling approach to regionalize and improve GloFAS-based river flow predictions. We used daily GloFAS-ERA5 data for the Sahzab, Mirkuh, and Markid flow stations in the Ajichai catchment, northwest Iran. A range of Machine Learning (ML) models were examined, including shallow learners, Feedforward Neural Network (FFNN), Adaptive Neuro-Fuzzy Inference System (ANFIS), Support Vector Regression (SVR), and the DL Long Short-Term Memory (LSTM) model. We tested these models and an ensemble of shallow learners under several scenarios that used both raw and bias-corrected GloFAS inputs. Precipitation and temperature from local weather stations and European Centre for Medium-Range Weather Forecasts (ECMWF) Reanalysis v5 (ERA5) were used as inputs. Observed discharge from 1993 until 2024 served as the target. In the multi-station strategy, upstream discharge from Sahzab and Mirkuh was included to improve downstream predictions at Markid. Integrating GloFAS data with advanced ML techniques, particularly ensemble learning within a multi-station framework, noticeably improved discharge prediction accuracy. Results, evaluated with Root Mean Square Error (RMSE) and the Coefficient of Determination (DC), show that the proposed Neural Averaging (NA) nonlinear ensemble outperforms individual ML models. At the Markid station, the multi-station approach improved performance over the single-station setup: RMSE decreased by ≈2.2% during calibration and ≈9.4% during verification, while DC increased by ≈3.9% and ≈25%, respectively. These improvements can support more reliable local flood forecasting and water-management decisions, and thus inform regional water policy and risk-reduction planning.
- Research Article
40
- 10.3390/app10093224
- May 6, 2020
- Applied Sciences
Monthly rainfall forecasts can be translated into monthly runoff predictions that could support water resources planning and management activities. Therefore, development of monthly rainfall forecasting models in reservoir watersheds is essential for generating future rainfall amounts as an input to a water-resources-system simulation model to predict water shortage conditions. This research aims to examine the reliability of linking a data preprocessing method (singular spectrum analysis, SSA) with machine learning, least-squares support vector regression (LS-SVR), and random forest (RF), for monthly rainfall forecasting in two reservoir watersheds (Deji and Shihmen reservoir watersheds) located in Taiwan. Merging SSA with LS-SVR and RF, the hybrid models (SSA-LSSVR and SSA-RF) were developed and compared with the standard models (LS-SVR and RF). The proposed models were calibrated and validated using the watersheds’ observed areal monthly rainfalls separated into 70 percent of data for calibration and 30 percent of data for validation. Model performances were evaluated using two accuracy measures, root mean square error (RMSE) and Nash–Sutcliffe efficiency (NSE). Results show that the hybrid models could efficiently forecast monthly rainfalls. Nonetheless, the performances of the hybrid models vary in both watersheds which suggests that prior knowledge about the watershed’s hydrological behavior would be helpful to implement the appropriate model. Overall, the hybrid models significantly surpass the standard models for the two studied watersheds, which indicates that the proposed models are a prudent modeling approach that could be employed in the current research regions for monthly rainfall forecasting.
- Research Article
84
- 10.3390/agronomy13051277
- Apr 28, 2023
- Agronomy
Timely and cost-effective crop yield prediction is vital in crop management decision-making. This study evaluates the efficacy of Unmanned Aerial Vehicle (UAV)-based Vegetation Indices (VIs) coupled with Machine Learning (ML) models for corn (Zea mays) yield prediction at vegetative (V6) and reproductive (R5) growth stages using a limited number of training samples at the farm scale. Four agronomic treatments, namely Austrian Winter Peas (AWP) (Pisum sativum L.) cover crop, biochar, gypsum, and fallow with sixteen replications were applied during the non-growing corn season to assess their impact on the following corn yield. Thirty different variables (i.e., four spectral bands: green, red, red edge, and near-infrared and twenty-six VIs) were derived from UAV multispectral data collected at the V6 and R5 stages to assess their utility in yield prediction. Five different ML algorithms including Linear Regression (LR), k-Nearest Neighbor (KNN), Random Forest (RF), Support Vector Regression (SVR), and Deep Neural Network (DNN) were evaluated in yield prediction. One-year experimental results of different treatments indicated a negligible impact on overall corn yield. Red edge, canopy chlorophyll content index, red edge chlorophyll index, chlorophyll absorption ratio index, green normalized difference vegetation index, green spectral band, and chlorophyll vegetation index were among the most suitable variables in predicting corn yield. The SVR predicted yield for the fallow with a Coefficient of Determination (R2) and Root Mean Square Error (RMSE) of 0.84 and 0.69 Mg/ha at V6 and 0.83 and 1.05 Mg/ha at the R5 stage, respectively. The KNN achieved a higher prediction accuracy for AWP (R2 = 0.69 and RMSE = 1.05 Mg/ha at V6 and 0.64 and 1.13 Mg/ha at R5) and gypsum treatment (R2 = 0.61 and RMSE = 1.49 Mg/ha at V6 and 0.80 and 1.35 Mg/ha at R5). The DNN achieved a higher prediction accuracy for biochar treatment (R2 = 0.71 and RMSE = 1.08 Mg/ha at V6 and 0.74 and 1.27 Mg/ha at R5). For the combined (AWP, biochar, gypsum, and fallow) treatment, the SVR produced the most accurate yield prediction with an R2 and RMSE of 0.36 and 1.48 Mg/ha at V6 and 0.41 and 1.43 Mg/ha at the R5. Overall, the treatment-specific yield prediction was more accurate than the combined treatment. Yield was most accurately predicted for fallow than other treatments regardless of the ML model used. SVR and KNN outperformed other ML models in yield prediction. Yields were predicted with similar accuracy at both growth stages. Thus, this study demonstrated that VIs coupled with ML models can be used in multi-stage corn yield prediction at the farm scale, even with a limited number of training data.
- Research Article
- 10.1016/j.mlwa.2026.100880
- Jun 1, 2026
- Machine Learning with Applications
Comparing allometric models to machine learning models for aboveground biomass estimation in agroforestry systems in Kenya
- Research Article
1
- 10.3390/jcm14186373
- Sep 10, 2025
- Journal of Clinical Medicine
Background: Suicide remains a leading cause of death among youth, yet effective tools to predict suicide attempts (SA) in individuals under 18 are scarce. This study aims to develop machine learning (ML) models to predict SA in paediatric populations using Google Trends data. Methods: Relative Search Volumes (RSVs) from Google Trends were analysed for terms linked to suicide risk factors. Pearson Correlation Coefficients (PCC) identified terms strongly associated with SA rates. Based on these, several ML models were developed and evaluated, including Random Forest Regression, Support Vector Regression (SVR), XGBoost, and Linear Regression. Model performance was assessed using metrics such as PCC, mean absolute error (MAE), mean squared error (MSE), root mean square error (RMSE), and mean absolute percentage error (MAPE). Results: Terms related to suicide prevention and symptoms, including psychiatrist and anxiety disorder, showed the strongest correlations with SA rates (PCC ≥ 0.90). Random Forest Regression emerged as the top-performing ML model (PCC = 0.953, MAPE = 20.12%, RMSE = 17.21), highlighting burnout, anxiety disorder, antidepressants, and psychiatrist as key predictors of SA. Other models’ scores were XGBoost (PCC = 0.446, MAPE = 22.57%, RMSE = 18.03), SVR (PCC = 0.833, MAPE = 42.23%, RMSE = 47.32) and Linear Regression (PCC = 0.947, MAPE = 23.64%, RMSE = 17.66). Conclusions: Google Trends–based ML models suggest potential utility for short-term prediction of youth SA. These preliminary findings support the utility of search data in identifying real-time suicide risk in paediatric populations.
- Research Article
26
- 10.3390/w14071032
- Mar 24, 2022
- Water
Stream temperature (Ts) is an important water quality parameter that affects ecosystem health and human water use for beneficial purposes. Accurate Ts predictions at different spatial and temporal scales can inform water management decisions that account for the effects of changing climate and extreme events. In particular, widespread predictions of Ts in unmonitored stream reaches can enable decision makers to be responsive to changes caused by unforeseen disturbances. In this study, we demonstrate the use of classical machine learning (ML) models, support vector regression and gradient boosted trees (XGBoost), for monthly Ts predictions in 78 pristine and human-impacted catchments of the Mid-Atlantic and Pacific Northwest hydrologic regions spanning different geologies, climate, and land use. The ML models were trained using long-term monitoring data from 1980–2020 for three scenarios: (1) temporal predictions at a single site, (2) temporal predictions for multiple sites within a region, and (3) spatiotemporal predictions in unmonitored basins (PUB). In the first two scenarios, the ML models predicted Ts with median root mean squared errors (RMSE) of 0.69–0.84 °C and 0.92–1.02 °C across different model types for the temporal predictions at single and multiple sites respectively. For the PUB scenario, we used a bootstrap aggregation approach using models trained with different subsets of data, for which an ensemble XGBoost implementation outperformed all other modeling configurations (median RMSE 0.62 °C).The ML models improved median monthly Ts estimates compared to baseline statistical multi-linear regression models by 15–48% depending on the site and scenario. Air temperature was found to be the primary driver of monthly Ts for all sites, with secondary influence of month of the year (seasonality) and solar radiation, while discharge was a significant predictor at only 10 sites. The predictive performance of the ML models was robust to configuration changes in model setup and inputs, but was influenced by the distance to the nearest dam with RMSE <1 °C at sites situated greater than 16 and 44 km from a dam for the temporal single site and regional scenarios, and over 1.4 km from a dam for the PUB scenario. Our results show that classical ML models with solely meteorological inputs can be used for spatial and temporal predictions of monthly Ts in pristine and managed basins with reasonable (<1 °C) accuracy for most locations.
- Research Article
- 10.3390/met16030266
- Feb 27, 2026
- Metals
Research on eco-friendly and energy-efficient machining processes has gained significant importance within the domain of sustainable production. This study is focused on enhancing the energy performance and sustainability of the milling process. Four machine learning (ML) models, namely, multiple linear regression (MLR), support vector regression (SVR), Gaussian process regression (GPR), and adaptive network-based fuzzy inference system (ANFIS), were proposed to estimate specific energy consumption (SEC) in the milling of Ti6-Al4-V under two eco-benign cooling conditions: cryogenic and minimum quantity lubrication (MQL). Several statistical metrics, including normalized mean absolute error (nMAE), mean absolute percentage error (MAPE), normalized root mean square error (nRMSE), maximum absolute percentage error (maxAPE), coefficient of determination (R2), and Willmott’s index of agreement (IA), were employed to validate the performances of the ML models. A high level of agreement between the predicted and experimental SEC data for both the training and test datasets supports the reliability of the proposed ML models. Although the MLR model performed well, the results revealed that the other ML models demonstrated better overall performance. According to the statistical metrics, the models’ predictive performance improved in the following sequence: MLR, SVR, GPR, and finally ANFIS, which demonstrated the highest predictive capability.
- Research Article
14
- 10.2166/wcc.2023.003
- Jun 20, 2023
- Journal of Water and Climate Change
A study was carried out to develop and evaluate the performance of different machine learning (ML) models for predicting reference evapotranspiration (ET0). The models included multiple linear regression (MLR), least square-support vector machine (LS-SVM), artificial neural networks (ANNs) and adaptive neuro-fuzzy inference system (ANFIS). The daily meteorological data for 50 years (1970–2019) were used to estimate ET0 using FAO-ET calculator. The FAO-ET calculator was compared with ML models to investigate the best-fit ML model for predicting ET. Thereafter, ET predicted by the best-fit ML model was compared with satellite (Moderate Resolution Imaging Spectroradiometer – MODIS) ET, which was finally mapped to a larger landscape (over entire Punjab and Haryana). Modeling of ET0 was best performed through LS-SVM followed by ANN2, ANN1, ANFIS10, ANFIS2, MLR and ANFIS9 models. Among developed models, coefficient of determination (R2) value varied from 0.800 to 0.998, being highest (0.998) under LS-SVM model. MODIS overestimated ET when compared with LS-SVM having R2 and root mean square error (RMSE) values of 0.73 and 3.95 mm, respectively. After applying the bias correction factor, R2 and RMSE were 0.74 and 1.19 mm, respectively. The ML and satellite-based ET estimation would be useful for timely water budgeting to manage the water scarcity problems from local to regional levels.
- Research Article
79
- 10.1371/journal.pone.0317619
- Jan 23, 2025
- PLOS ONE
This study presents a comprehensive comparative analysis of Machine Learning (ML) and Deep Learning (DL) models for predicting Wind Turbine (WT) power output based on environmental variables such as temperature, humidity, wind speed, and wind direction. Along with Artificial Neural Network (ANN), Long Short-Term Memory (LSTM), Recurrent Neural Network (RNN), and Convolutional Neural Network (CNN), the following ML models were looked at: Linear Regression (LR), Support Vector Regressor (SVR), Random Forest (RF), Extra Trees (ET), Adaptive Boosting (AdaBoost), Categorical Boosting (CatBoost), Extreme Gradient Boosting (XGBoost), and Light Gradient Boosting Machine (LightGBM). Using a dataset of 40,000 observations, the models were assessed based on R-squared, Mean Absolute Error (MAE), and Root Mean Square Error (RMSE). ET achieved the highest performance among ML models, with an R-squared value of 0.7231 and a RMSE of 0.1512. Among DL models, ANN demonstrated the best performance, achieving an R-squared value of 0.7248 and a RMSE of 0.1516. The results show that DL models, especially ANN, did slightly better than the best ML models. This means that they are better at modeling non-linear dependencies in multivariate data. Preprocessing techniques, including feature scaling and parameter tuning, improved model performance by enhancing data consistency and optimizing hyperparameters. When compared to previous benchmarks, the performance of both ANN and ET demonstrates significant predictive accuracy gains in WT power output forecasting. This study’s novelty lies in directly comparing a diverse range of ML and DL algorithms while highlighting the potential of advanced computational approaches for renewable energy optimization.
- Research Article
40
- 10.2196/47833
- Nov 20, 2023
- JMIR Medical Informatics
Machine learning (ML) models provide more choices to patients with diabetes mellitus (DM) to more properly manage blood glucose (BG) levels. However, because of numerous types of ML algorithms, choosing an appropriate model is vitally important. In a systematic review and network meta-analysis, this study aimed to comprehensively assess the performance of ML models in predicting BG levels. In addition, we assessed ML models used to detect and predict adverse BG (hypoglycemia) events by calculating pooled estimates of sensitivity and specificity. PubMed, Embase, Web of Science, and Institute of Electrical and Electronics Engineers Explore databases were systematically searched for studies on predicting BG levels and predicting or detecting adverse BG events using ML models, from inception to November 2022. Studies that assessed the performance of different ML models in predicting or detecting BG levels or adverse BG events of patients with DM were included. Studies with no derivation or performance metrics of ML models were excluded. The Quality Assessment of Diagnostic Accuracy Studies tool was applied to assess the quality of included studies. Primary outcomes were the relative ranking of ML models for predicting BG levels in different prediction horizons (PHs) and pooled estimates of the sensitivity and specificity of ML models in detecting or predicting adverse BG events. In total, 46 eligible studies were included for meta-analysis. Regarding ML models for predicting BG levels, the means of the absolute root mean square error (RMSE) in a PH of 15, 30, 45, and 60 minutes were 18.88 (SD 19.71), 21.40 (SD 12.56), 21.27 (SD 5.17), and 30.01 (SD 7.23) mg/dL, respectively. The neural network model (NNM) showed the highest relative performance in different PHs. Furthermore, the pooled estimates of the positive likelihood ratio and the negative likelihood ratio of ML models were 8.3 (95% CI 5.7-12.0) and 0.31 (95% CI 0.22-0.44), respectively, for predicting hypoglycemia and 2.4 (95% CI 1.6-3.7) and 0.37 (95% CI 0.29-0.46), respectively, for detecting hypoglycemia. Statistically significant high heterogeneity was detected in all subgroups, with different sources of heterogeneity. For predicting precise BG levels, the RMSE increases with a rise in the PH, and the NNM shows the highest relative performance among all the ML models. Meanwhile, current ML models have sufficient ability to predict adverse BG events, while their ability to detect adverse BG events needs to be enhanced. PROSPERO CRD42022375250; https://www.crd.york.ac.uk/prospero/display_record.php?RecordID=375250.
- Research Article
62
- 10.3390/su14042341
- Feb 18, 2022
- Sustainability
Over the last years, the global application of machine learning (ML) models in groundwater quality studies has proved to be a robust alternative tool to produce highly accurate results at a low cost. This research aims to evaluate the ability of machine learning (ML) models to predict the quality of groundwater for irrigation purposes in the downstream Medjerda river basin (DMB) in Tunisia. The random forest (RF), support vector regression (SVR), artificial neural networks (ANN), and adaptive boosting (AdaBoost) models were tested to predict the irrigation quality water parameters (IWQ): total dissolved solids (TDS), potential salinity (PS), sodium adsorption ratio (SAR), exchangeable sodium percentage (ESP), and magnesium adsorption ratio (MAR) through low-cost, in situ physicochemical parameters (T, pH, EC) as input variables. In view of this, seventy-two (72) representative groundwater samples have been collected and analysed for major cations and anions during pre-and post-monsoon seasons of 3 years (2019–2021) to compute IWQ parameters. The performance of the ML models was evaluated according to Pearson’s correlation coefficient (r), the root means square error (RMSE), and the relative bias (RBIAS). The model sensitivity analysis was evaluated to identify input parameters that considerably impact the model predictions using the one-factor-at-time (OFAT) method of the Monte Carlo (MC) approach. The results show that the AdaBoost model is the most appropriate model for predicting all parameters (r was ranged between 0.88 and 0.89), while the random forest model is suitable for predicting only four parameters: TDS, PS, SAR, and ESP (r was with 0.65 to 0.87). Added to that, this study found out that the ANN and SVR models perform well in predicting three parameters (TDS, PS, SAR) and two parameters (PS, SAR), respectively, with the most optimal value of generalization ability (GA) close to unity (between 1 and 0.98). Moreover, the results of the uncertainty analysis confirmed the prominent superiority and robustness of the ML models to produce excellent predictions with only a few physicochemical parameters as inputs. The developed ML models are relevant for predicting cost-effective irrigation water quality indices and can be applied as a DSS tool to improve water management in the Medjerda basin.
- Research Article
6
- 10.1016/j.cmpb.2025.108657
- Apr 1, 2025
- Computer methods and programs in biomedicine
Accurate estimation of resting energy expenditure (REE) is critical for guiding nutritional therapy in critically ill patients. While indirect calorimetry (IC) is the gold standard for REE measurement, it is not routinely feasible in clinical settings due to its complexity and cost. Predictive equations (PEs) offer a simpler alternative but are often inaccurate in critically ill populations. While recent advancements in machine learning (ML) and deep learning (DL) offer potential for improving REE estimation by capturing complex relationships between physiological variables, these approaches have not yet been widely applied or validated in critically ill populations. This prospective study compared the performance of nine commonly used PEs, including the Harris-Benedict (H-B1919), Penn State, and TAH equations, with ML models (XGBoost, Random Forest Regressor [RFR], Support Vector Regression), and DL models (Convolutional Neural Networks [CNN]) in estimating REE in critically ill patients. A dataset of 300 IC measurements from an intensive care unit (ICU) was used, with REE measured by both IC and PEs. The ML/DL models were trained using a combination of static (i.e., age, height, body weight) and dynamic (i.e., minute ventilation, body temperature) variables. A five-fold cross validation was performed to assess the model prediction performance using the root mean square error (RMSE) metric. Of the PEs analysed, H-B1919 yielded the lowest RMSE at 362 calories. However, the XGBoost and RFR models significantly outperformed all PEs, achieving RMSE values of 199 and 200 calories, respectively. The CNN model demonstrated the poorest performance among ML models, with an RMSE of 250 calories. The inclusion of additional categorical variables such as body mass index (BMI) and body temperature classes slightly reduced RMSE across ML and DL models. Despite data augmentation and imputation techniques, no significant improvements in model performance were observed. ML models, particularly XGBoost and RFR, provide more accurate REE estimations than traditional PEs, highlighting their potential to better capture the complex, non-linear relationships between physiological variables and REE. These models offer a promising alternative for guiding nutritional therapy in clinical settings, though further validation on independent datasets and across diverse patient populations is warranted.
- Research Article
30
- 10.3390/ma15134436
- Jun 23, 2022
- Materials
Bacterial-based self-healing concrete (BSHC) is a well-known healing technology which has been investigated for a few decades for its excellent crack healing capacity. Nevertheless, considered as costly and time-consuming, the healing performance (HP) of concrete with various types of bacteria can be designed and evaluated only in laboratory environments. Employing machine learning (ML) models for predicting the HP of BSHC is inspired by practical applications using concrete mechanical properties. The HP of BSHC can be predicted to save the time and cost of laboratory tests, bacteria selection and healing mechanisms adoption. In this paper, three types of BSHC, including ureolytic bacterial healing concrete (UBHC), aerobic bacterial healing concrete (ABHC) and nitrifying bacterial healing concrete (NBHC), and ML models with five kinds of algorithms consisting of the support vector regression (SVR), decision tree regression (DTR), deep neural network (DNN), gradient boosting regression (GBR) and random forest (RF) are established. Most importantly, 22 influencing factors are first employed as variables in the ML models to predict the HP of BSHC. A total of 797 sets of BSHC tests available in the open literature between 2000 and 2021 are collected to verify the ML models. The grid search algorithm (GSA) is also utilised for tuning parameters of the algorithms. Moreover, the coefficient of determination (R2) and root mean square error (RMSE) are applied to evaluate the prediction ability, including the prediction performance and accuracy of the ML models. The results exhibit that the GBR model has better prediction ability (R2GBR = 0.956, RMSEGBR = 6.756%) than other ML models. Finally, the influence of the variables on the HP is investigated by employing the sensitivity analysis in the GBR model.
- Research Article
9
- 10.13031/jnrae.15647
- Jan 1, 2023
- Journal of Natural Resources and Agricultural Ecosystems
Highlights Machine Learning (ML) models are identified, reviewed, and analyzed for HAB predictions. Data preprocessing is vital for efficient ML model development. ML models for toxin production and monitoring are limited. Abstract. Harmful algal blooms (HABs) are detrimental to livestock, humans, pets, the environment, and the global economy, which calls for a robust approach to their management. While process-based models can inform practitioners about HAB enabling conditions, they have inherent limitations in accurately predicting harmful algal blooms. To address these limitations, Machine Learning (ML) models can potentially leverage large volumes of IoT data to aid in near real-time predictions. ML models have evolved as efficient tools for understanding patterns and relationships between water quality parameters and HAB expansion. This review describes ML models currently used for predicting and forecasting HABs in freshwater ecosystems and presents model structures and their application for predicting algal parameters and related toxins. The review revealed that regression trees, random forest, Artificial Neural Network (ANN), Support Vector Regression (SVR), Long Short-Term Memory (LSTM), and Gated Recurrent Unit (GRU) are the most frequently used models for HABs monitoring. This review shows ML models' prowess in identifying significant variables influencing algal growth, HAB drivers, and multistep HAB prediction. Hybrid models also improve the prediction of algal-related parameters through improved optimization techniques and variable selection algorithms. While ML models often focus on algal biomass prediction, few studies apply ML models for toxin monitoring and prediction. This limitation can be associated with a lack of high-frequency toxin datasets for model development, and exploring this domain is encouraged. This review serves as a guide for policymakers and researchers to implement ML models for HAB prediction and reveals the potential of ML models for decision support and early prediction for HAB management. Keywords: Cyanobacteria, Freshwater, Harmful algal blooms, Machine learning, Water quality.