Accelerate Literature Icon
Want to do a literature review? Try our new Literature Review workflow

Salinity Forecasting in the Vietnamese Mekong Delta: Evaluating the Predictive Power of Machine Learning Approaches using Multitemporal Lag Features

  • Abstract
  • Literature Map
  • Similar Papers
Abstract
Translate article icon Translate Article Star icon

Salinity intrusion threatens water security, agriculture, and livelihoods in the Vietnamese Mekong Delta (VMD), particularly under the combined pressures of climate change and upstream hydrological developments. Accurate short- to mid-term salinity forecasts are essential for proactive water resource management. This study evaluates the performance of two machine learning models—Random Forest Regression (RFR) and Support Vector Regression (SVR)—for salinity forecasting using long-term observational data (1996–2023) from 44 monitoring stations in the VMD. Multitemporal lag features (1, 10, 20, 30, and 60 days) were generated from observed salinity records to capture temporal dynamics. Bayesian optimization and time-series cross-validation were used for model tuning. Results show that SVR performs best for short-term forecasts (1–3 days), while RFR provides more stable predictions over longer horizons (4–7 days). Lag windows of 20–30 days yielded the most accurate results, likely reflecting tidal cycle influences. These findings demonstrate the potential of machine learning approaches to improve salinity early warning systems and support adaptive management in deltaic regions affected by saltwater intrusion.

Similar Papers
  • PDF Download Icon
  • Research Article
  • Cite Count Icon 71
  • 10.5194/bg-14-5551-2017
Empirical methods for the estimation of Southern Ocean CO 2 : support vector and random forest regression
  • Dec 8, 2017
  • Biogeosciences
  • Luke Gregor + 2 more

Abstract. The Southern Ocean accounts for 40 % of oceanic CO2 uptake, but the estimates are bound by large uncertainties due to a paucity in observations. Gap-filling empirical methods have been used to good effect to approximate pCO2 from satellite observable variables in other parts of the ocean, but many of these methods are not in agreement in the Southern Ocean. In this study we propose two additional methods that perform well in the Southern Ocean: support vector regression (SVR) and random forest regression (RFR). The methods are used to estimate ΔpCO2 in the Southern Ocean based on SOCAT v3, achieving similar trends to the SOM-FFN method by Landschützer et al. (2014). Results show that the SOM-FFN and RFR approaches have RMSEs of similar magnitude (14.84 and 16.45 µatm, where 1 atm = 101 325 Pa) where the SVR method has a larger RMSE (24.40 µatm). However, the larger errors for SVR and RFR are, in part, due to an increase in coastal observations from SOCAT v2 to v3, where the SOM-FFN method used v2 data. The success of both SOM-FFN and RFR depends on the ability to adapt to different modes of variability. The SOM-FFN achieves this by having independent regression models for each cluster, while this flexibility is intrinsic to the RFR method. Analyses of the estimates shows that the SVR and RFR's respective sensitivity and robustness to outliers define the outcome significantly. Further analyses on the methods were performed by using a synthetic dataset to assess the following: which method (RFR or SVR) has the best performance? What is the effect of using time, latitude and longitude as proxy variables on ΔpCO2? What is the impact of the sampling bias in the SOCAT v3 dataset on the estimates? We find that while RFR is indeed better than SVR, the ensemble of the two methods outperforms either one, due to complementary strengths and weaknesses of the methods. Results also show that for the RFR and SVR implementations, it is better to include coordinates as proxy variables as RMSE scores are lowered and the phasing of the seasonal cycle is more accurate. Lastly, we show that there is only a weak bias due to undersampling. The synthetic data provide a useful framework to test methods in regions of sparse data coverage and show potential as a useful tool to evaluate methods in future studies.

  • PDF Download Icon
  • Research Article
  • Cite Count Icon 37
  • 10.3390/w13070920
Prediction of River Stage Using Multistep-Ahead Machine Learning Techniques for a Tidal River of Taiwan
  • Mar 27, 2021
  • Water
  • Wen-Dar Guo + 4 more

Time-series prediction of a river stage during typhoons or storms is essential for flood control or flood disaster prevention. Data-driven models using machine learning (ML) techniques have become an attractive and effective approach to modeling and analyzing river stage dynamics. However, relatively new ML techniques, such as the light gradient boosting machine regression (LGBMR), have rarely been applied to predict the river stage in a tidal river. In this study, data-driven ML models were developed under a multistep-ahead prediction framework and evaluated for river stage modeling. Four ML techniques, namely support vector regression (SVR), random forest regression (RFR), multilayer perceptron regression (MLPR), and LGBMR, were employed to establish data-driven ML models with Bayesian optimization. The models were applied to simulate river stage hydrographs of the tidal reach of the Lan-Yang River Basin in Northeastern Taiwan. Historical measurements of rainfall, river stages, and tidal levels were collected from 2004 to 2017 and used for training and validation of the four models. Four scenarios were used to investigate the effect of the combinations of input variables on river stage predictions. The results indicated that (1) the tidal level at a previous stage significantly affected the prediction results; (2) the LGBMR model achieves more favorable prediction performance than the SVR, RFR, and MLPR models; and (3) the LGBMR model could efficiently and accurately predict the 1–6-h river stage in the tidal river. This study provides an extensive and insightful comparison of four data-driven ML models for river stage forecasting that can be helpful for model selection and flood mitigation.

  • Research Article
  • Cite Count Icon 5
  • 10.1007/s13201-025-02419-z
Short-term salinity prediction for coastal areas of the Vietnamese Mekong Delta using various machine learning algorithms: a case study in Soc Trang Province
  • Mar 19, 2025
  • Applied Water Science
  • Le Thi Thanh Dang + 4 more

Saltwater intrusion has significant and diverse impacts on agriculture, freshwater resources, and the well-being of coastal communities. To effectively address this issue, precise models for predicting saltwater intrusion must be developed, as well as timely information for reaction planning. In this study, a spectrum of machine learning (ML) methodologies, specifically Random Forest Regression (RFR), Support Vector Regression (SVR), Long Short-Term Memory (LSTM), Artificial Neural Network (ANN), Extreme Gradient Boosting (XGBoost), and Ridge Regression (RR), was systematically employed to predict salinity levels within the coastal environs of the Mekong Delta, Vietnam. The input dataset comprised hourly salinity measurements from Tran De, Long Phu, Dai Ngai, and Soc Trang stations and hourly water-level data from Tran De station and hourly discharge data from the Can Tho hydrological station. The dataset was partitioned into two distinct sets for the purpose of model development and evaluation, employing a division ratio of 75% for training (constituting 8469 observations) and 25% for testing (comprising 2822 observations). The results indicate that ML models are suitable for short-term salinity prediction, with a forecasting time of up to 16 h in this area. These research findings highlight the potential of machine learning in addressing saltwater intrusion and provide valuable insights for developing appropriate response policies. By leveraging the strengths of these models and considering the optimal forecasting time, policymakers can make informed decisions and implement effective measures to mitigate the impacts of saltwater intrusion in the Mekong Delta.

  • Research Article
  • Cite Count Icon 5
  • 10.1016/j.geits.2025.100303
Data-driven machine learning techniques for fuel economy prediction in sustainable transportation systems
  • Feb 1, 2026
  • Green Energy and Intelligent Transportation
  • Muhammad Sohaib Zahid + 1 more

Data-driven machine learning techniques for fuel economy prediction in sustainable transportation systems

  • Research Article
  • 10.53560/ppasa(60-4)820
Measuring the Performance of Supervised Machine Learning Algorithms for Optimizing Wheat Productivity Prediction Models: A Comparative Study
  • Dec 12, 2023
  • Proceedings of the Pakistan Academy of Sciences: A. Physical and Computational Sciences
  • Malik Muhammad Hussain + 5 more

The issue of precise crop prediction gained worldwide attention in the midst of food security concerns. In this study, the efficacies of different machine learning (ML) algorithms, i.e., multiple linear regression (MLR), decision tree regression (DTR), random forest regression (RFR), and support vector regression (SVR) are integrated to predict wheat productivity. The performances of ML algorithms are then measured to get the optimized model. The updated dataset is collected from the Crop Reporting Service for various agronomical constraints. Randomized data partitions, hyper-parametric tuning, complexity analysis, cross-validation measures, learning curves, evaluation metrics and prediction errors are used to get the optimized model. ML model is applied using 75% training dataset and 25% testing datasets. RFR achieved the highest R2 value of 0.90 for the training model, followed by DTR, MLR, and SVR. In the testing model, RFR also achieved an R2 value of 0.74, followed by MLR, DTR, and SVR. The lowest prediction error (P.E) is found for the RFR, followed by DTR, MLR, and SVR. K-Fold cross-validation measures also depict that RFR is an optimized model when compared with DTR, MLR and SVR.

  • PDF Download Icon
  • Research Article
  • Cite Count Icon 9
  • 10.3389/fpls.2022.821365
Poplar\u2019s Waterlogging Resistance Modeling and Evaluating: Exploring and Perfecting the Feasibility of Machine Learning Methods in Plant Science
  • Feb 11, 2022
  • Frontiers in Plant Science
  • Xuelin Xie + 3 more

Floods, as one of the most common disasters in the natural environment, have caused huge losses to human life and property. Predicting the flood resistance of poplar can effectively help researchers select seedlings scientifically and resist floods precisely. Using machine learning algorithms, models of poplar’s waterlogging tolerance were established and evaluated. First of all, the evaluation indexes of poplar’s waterlogging tolerance were analyzed and determined. Then, significance testing, correlation analysis, and three feature selection algorithms (Hierarchical clustering, Lasso, and Stepwise regression) were used to screen photosynthesis, chlorophyll fluorescence, and environmental parameters. Based on this, four machine learning methods, BP neural network regression (BPR), extreme learning machine regression (ELMR), support vector regression (SVR), and random forest regression (RFR) were used to predict the flood resistance of poplar. The results show that random forest regression (RFR) and support vector regression (SVR) have high precision. On the test set, the coefficient of determination (R2) is 0.8351 and 0.6864, the root mean square error (RMSE) is 0.2016 and 0.2780, and the mean absolute error (MAE) is 0.1782 and 0.2031, respectively. Therefore, random forest regression (RFR) and support vector regression (SVR) can be given priority to predict poplar flood resistance.

  • PDF Download Icon
  • Research Article
  • Cite Count Icon 19
  • 10.1007/s00170-024-14087-5
Effects of process parameters on the surface characteristics of laser powder bed fusion printed parts: machine learning predictions with random forest and support vector regression
  • Jul 5, 2024
  • The International Journal of Advanced Manufacturing Technology
  • Naol Dessalegn Dejene + 2 more

Laser powder bed fusion (L-PBF) fuses metallic powder using a high-energy laser beam, forming parts layer by layer. This technique offers flexibility and design freedom in metal additive manufacturing (MAM). However, achieving the desired surface quality remains challenging and impacts functionality and reliability. L-PBF process parameters significantly influence surface roughness. Identifying the most critical factors among numerous parameters is essential for improving quality. This study examines the effects of key process parameters on the surface roughness of AlSi10Mg, a widely used aluminum alloy in high-tech industries, fabricated by L-PBF. Part orientation, laser power, scanning speed, and layer thickness were identified as crucial parameters via cause-and-effect analysis. To systematically examine their effects, the Taguchi method was employed within the framework of the design of experiment (DoE). Experimental results and statistical analysis revealed that laser power, scanning speed, and layer thickness significantly influence surface roughness parameters: arithmetic mean (Ra) and root mean square (Rq). Main effect plots and energy density analyses confirmed their impact on surface quality. Microscopic investigations identified surface flaws such as spattering, balling, and porosity contributing to poor quality. Given the complex interplay between parameters and surface quality, accurately predicting their effects is challenging. To address this, machine learning models, specifically random forest regression (RFR) and support vector regression (SVR), were used to predict the effects on surface roughness. The RFR model’s R2 values for predicting Ra and Rq are 97% and 85%, while the SVR model’s predictions are 85% and 66%, respectively. Evaluation metrics demonstrated that the RFR model outperformed SVR in predicting surface roughness.

  • Research Article
  • Cite Count Icon 22
  • 10.1016/j.procs.2020.11.031
COVID-19 in Bangladesh: A Deeper Outlook into The Forecast with Prediction of Upcoming Per Day Cases Using Time Series.
  • Jan 1, 2020
  • Procedia Computer Science
  • Abu Kaisar Mohammad Masum + 4 more

COVID-19 in Bangladesh: A Deeper Outlook into The Forecast with Prediction of Upcoming Per Day Cases Using Time Series.

  • PDF Download Icon
  • Research Article
  • Cite Count Icon 25
  • 10.1038/s41598-024-66699-2
Mapping reservoir water quality from Sentinel-2 satellite data based on a new approach of weighted averaging: Application of Bayesian maximum entropy
  • Jul 16, 2024
  • Scientific Reports
  • Mohammad Reza Nikoo + 5 more

In regions like Oman, which are characterized by aridity, enhancing the water quality discharged from reservoirs poses considerable challenges. This predicament is notably pronounced at Wadi Dayqah Dam (WDD), where meeting the demand for ample, superior water downstream proves to be a formidable task. Thus, accurately estimating and mapping water quality indicators (WQIs) is paramount for sustainable planning of inland in the study area. Since traditional procedures to collect water quality data are time-consuming, labor-intensive, and costly, water resources management has shifted from gathering field measurement data to utilizing remote sensing (RS) data. WDD has been threatened by various driving forces in recent years, such as contamination from different sources, sedimentation, nutrient runoff, salinity intrusion, temperature fluctuations, and microbial contamination. Therefore, this study aimed to retrieve and map WQIs, namely dissolved oxygen (DO) and chlorophyll-a (Chl-a) of the Wadi Dayqah Dam (WDD) reservoir from Sentinel-2 (S2) satellite data using a new procedure of weighted averaging, namely Bayesian Maximum Entropy-based Fusion (BMEF). To do so, the outputs of four Machine Learning (ML) algorithms, namely Multilayer Regression (MLR), Random Forest Regression (RFR), Support Vector Regression (SVRs), and XGBoost, were combined using this approach together, considering uncertainty. Water samples from 254 systematic plots were obtained for temperature (T), electrical conductivity (EC), chlorophyll-a (Chl-a), pH, oxidation–reduction potential (ORP), and dissolved oxygen (DO) in WDD. The findings indicated that, throughout both the training and testing phases, the BMEF model outperformed individual machine learning models. Considering Chl-a, as WQI, and R-squared, as evaluation indices, BMEF outperformed MLR, SVR, RFR, and XGBoost by 6%, 9%, 2%, and 7%, respectively. Furthermore, the results were significantly enhanced when the best combination of various spectral bands was considered to estimate specific WQIs instead of using all S2 bands as input variables of the ML algorithms.

  • Research Article
  • Cite Count Icon 25
  • 10.1016/j.jhydrol.2024.130655
Extent of saltwater intrusion and freshwater exploitability in the coastal Vietnamese Mekong Delta assessed by gauging records and numerical simulations
  • Jan 19, 2024
  • Journal of Hydrology
  • Dung Duc Tran + 5 more

Extent of saltwater intrusion and freshwater exploitability in the coastal Vietnamese Mekong Delta assessed by gauging records and numerical simulations

  • Research Article
  • Cite Count Icon 9
  • 10.9734/air/2020/v21i930238
Assessment of the Different Machine Learning Models for Prediction of Cluster Bean (Cyamopsis tetragonoloba L. Taub.) Yield
  • Aug 27, 2020
  • Advances in Research
  • Darshan Jagannath Pangarkar + 3 more

Prediction of crop yield can help traders, agri-business and government agencies to plan their activities accordingly. It can help government agencies to manage situations like over or under production. Traditionally statistical and crop simulation methods are used for this purpose. Machine learning models can be great deal of help. Aim of present study is to assess the predictive ability of various machine learning models for Cluster bean (Cyamopsis tetragonoloba L. Taub.) yield prediction. Various machine learning models were applied and tested on panel data of 19 years i.e. from 1999-2000 to 2017-18 for the Bikaner district of Rajasthan. Various data mining steps were performed before building a model. K- Nearest Nighbors (K-NN), Support Vector Regression (SVR) with various kernels, and Random forest regression were applied. Cross validation was also performed to know extra sampler validity. The best fitted model was chosen based cross validation scores and R2 values. Besides the coefficient of determination (R2), root mean squared error (RMSE), mean absolute error (MAE), and root relative squared error (RRSE) were calculated for the testing set. Support vector regression with linear kernel has the lowest RMSE (23.19), RRSE (0.14), MAE (19.27) values followed by random forest regression and second-degree polynomial support vector regression with the value of gamma = auto. Instead there was a little difference with R2, placing support vector regression first (98.31%), followed by second-degree polynomial support vector regression with value of gamma = auto (89.83%) and second-degree polynomial support vector regression with value of gamma = scale (88.83%). On two-fold cross validation, support vector regression with a linear kernel had the highest cross validation score explaining 71% (+/-0.03) followed by second-degree polynomial support vector regression with a value of gamma = auto and random forest regression. KNN and support vector regression with radial basis function as a kernel function had negative cross validation scores. Support vector regression with linear kernel was found to be the best-fitted model for predicting the yield as it had higher sample validity (98.31%) and global validity (71%).

  • PDF Download Icon
  • Research Article
  • Cite Count Icon 24
  • 10.3390/rs15092264
Machine Learning Algorithms for the Retrieval of Canopy Chlorophyll Content and Leaf Area Index of Crops Using the PROSAIL-D Model with the Adjusted Average Leaf Angle
  • Apr 25, 2023
  • Remote Sensing
  • Qi Sun + 5 more

The canopy chlorophyll content (CCC) and leaf area index (LAI) are both essential indicators for crop growth monitoring and yield estimation. The PROSAIL model, which couples the properties optique spectrales des feuilles (PROSPECT) and scattering by arbitrarily inclined leaves (SAIL) radiative transfer models, is commonly used for the quantitative retrieval of crop parameters; however, its homogeneous canopy assumption limits its accuracy, especially in the case of multiple crop categories. The adjusted average leaf angle (ALAadj), which can be parameterized for a specific crop type, increases the applicability of the PROSAIL model for specific crop types with a non-uniform canopy and has the potential to enhance the performance of PROSAIL-coupled hybrid methods. In this study, the PROSAIL-D model was used to generate the ALAadj values of wheat, soybean, and maize crops based on ground-measured spectra, the LAI, and the leaf chlorophyll content (LCC). The results revealed ALAadj values of 62 degrees for wheat, 45 degrees for soybean, and 60 degrees for maize. Support vector regression (SVR), random forest regression (RFR), extremely randomized trees regression (ETR), the gradient boosting regression tree (GBRT), and stacking learning (STL) were applied to simulated data of the ALAadj in 50-band data to retrieve the CCC and LAI of the crops. The results demonstrated that the estimation accuracy of singular crop parameters, particularly the crop LAI, was greatly enhanced by the five machine learning methods on the basis of data simulated with the ALAadj. Regarding the estimation results of mixed crops, the machine learning algorithms using ALAadj datasets resulted in estimations of CCC (RMSE: RFR = 51.1 μg cm−2, ETR = 54.7 μg cm−2, GBRT = 54.9 μg cm−2, STL = 48.3 μg cm−2) and LAI (RMSE: SVR = 0.91, RFR = 1.03, ETR = 1.05, GBRT = 1.05, STL = 0.97), that outperformed the estimations without using the ALAadj (namely CCC RMSE: RFR = 93.0 μg cm−2, ETR = 60.1 μg cm−2, GBRT = 60.0 μg cm−2, STL = 68.5 μg cm−2 and LAI RMSE: SVR = 2.10, RFR = 2.28, ETR = 1.67, GBRT = 1.66, STL = 1.51). Similar findings were obtained using the suggested method in conjunction with 19-band data, demonstrating the promising potential of this method to estimate the CCC and LAI of crops at the satellite scale.

  • Research Article
  • Cite Count Icon 43
  • 10.1016/j.jenvman.2023.118782
Improving groundwater nitrate concentration prediction using local ensemble of machine learning models
  • Aug 17, 2023
  • Journal of Environmental Management
  • Hojjatollah Mahboobi + 2 more

Improving groundwater nitrate concentration prediction using local ensemble of machine learning models

  • Research Article
  • Cite Count Icon 2
  • 10.1038/s41598-025-24926-4
A novel approach of developing machine learning based models for the prediction of facial dimensions from dental parameters
  • Nov 20, 2025
  • Scientific Reports
  • Damini Siwan + 3 more

Personal identification of an individual has always been a major concern in forensic science. Reconstruction of the facial profile is considered as one of the final stages in the process of identification. Nevertheless, recent advancements in artificial intelligence (AI) and machine learning (ML) have demonstrated remarkable potential in predictive modelling and forensic applications. The current study uses customised machine learning models to predict facial dimensions based on dental and jaw parameters. A sample of 422 participants (201 males and 221 females) from a North Indian population was collected and analysed. Dental casts, anthropometric facial measurements and photographs of the participants were collected with informed consent. ML models such as Support Vector Regression (SVR), Random Forest Regression (RFR), Decision Tree Regression (DTR), and Linear Regression (LR) were trained using dental and jaw measurements as input features for the models. The results show that the ML models predicted the facial dimensions with an accuracy of 90–94% and a very low prediction error of 0.1–0.9 across all facial measurements. Among the models, SVR and LR models perform well, followed by RFR, whereas DFR yielded comparatively lower accuracy. The findings demonstrate that machine learning models (SVR, RFR, DTR, and LR) can be used as novel approach to predict facial dimensions from jaw and teeth parameters. These techniques can be combined with other facial reconstruction techniques to produce more precise and accurate outcomes. The reliability and accuracy in predicting the facial dimensions indicate that the results can be applied in the practical and real situations such as personal identification, forensic investigations, disaster victim identification cases, and archaeological remains where only jaw and teeth are available for examination. Integrating ML-based predictions with traditional facial reconstruction techniques could enhance the accuracy and reliability of forensic identification methodologies.

  • PDF Download Icon
  • Research Article
  • Cite Count Icon 4
  • 10.1038/s41598-023-46171-3
Machine learning models to predict post-dialysis blood pressure in children and young adults on maintenance hemodialysis
  • Nov 4, 2023
  • Scientific Reports
  • Raed Bou-Matar + 2 more

Hypertension is associated with significant cardiovascular morbidity. Blood pressure (BP) control on maintenance hemodialysis (HD) is strongly impacted by volume status. The objective of this study was to assess whether machine learning (ML) is effective in predicting post-HD BP in children and young adults on HD. We collected data on BP, IDWG, pulse, and weights for patients on maintenance HD (> 3 months). Input features included DW, pre-post weight difference, IDWG and pre-HD BP. Seven models were trained and tuned using open-source libraries. Model performance was evaluated using time-series cross-validation on a rolling basis (3–12 month training, 1-day testing). Various regression scores were compared between models. Data for 35 patients (14,375 HD sessions) were analyzed. Extreme gradient boosting (XGB) and vector autoregression with exogenous regressors (VARX) achieved better accuracy in predicting post-dialysis systolic BP than K-nearest neighbor, support vector regression (SVR) with radial basis function kernel and random forest (p < 0.001 for each). The differences in accuracy between XGB, VARX, SVR with linear kernel, random forest and linear regression were not significant. Using clinical parameters, ML models may be useful in predicting post-HD BP, which may help guide DW adjustment and optimizing BP control for maintenance HD patients.

Save Icon
Up Arrow
Open/Close
Notes

Save Important notes in documents

Highlight text to save as a note, or write notes directly

You can also access these Documents in Paperpal, our AI writing tool

Powered by our AI Writing Assistant