Health effects due to solid biomass cooking emissions: A survey-based study in Haryana and Rajasthan
• Surveyed 1,000 rural households in Mahendragarh (Haryana) and Jhunjhunu (Rajasthan) to assess health impacts of biomass fuel use. • Identified high prevalence of respiratory symptoms, especially among women and children exposed to smoke from wood, dung cakes, and crop residues. • Developed a comprehensive analytical system integrating PM monitoring, weather data, and health surveys to model exposure-risk relationships. • Applied machine learning (Random Forest, XGBoost, LSTM) and causal inference methods (Granger causality, Causal Impact) to predict pollution and health outcomes. • Designed real-time API and dashboard tools for public health alerts and policy support in rural air quality management. This study investigates the spatiotemporal dynamics of ambient air pollution and its health impacts across two semi-urban districts in India- Jhunjhunu and Mahendragarh, using a multidisciplinary approach combining statistical analysis, machine learning, and causal inference. A one-year high-resolution monitoring dataset of PM₁, PM₂.₅, PM₄, and PM₁₀ was integrated with structured household health surveys covering over 1,000 households. High-resolution monitoring of PM₁, PM 2.5 , PM₄, and PM₁₀, along with survey-based health data, was analyzed to explore pollutant behavior, exposure-response relationships, and symptom prevalence. Linear regression models effectively predicted PM 2.5 trends in Jhunjhunu, while advanced models such as Random Forest, XGBoost, and Long Short-Term Memory (LSTM) captured complex variability in Mahendragarh. Models were trained using a 70:30 train–test split with k-fold cross-validation and evaluated using RMSE, MAE, and R² metrics. LSTM and XGBoost achieved the best performance (R² up to 0.87; RMSE reduced by approximately 30% compared to linear regression). SHAP analysis highlighted PM₁ and PM₄ as critical predictors, underscoring the need to expand national air quality standards beyond PM 2.5 and PM₁₀. Explainable machine learning using SHAP identified PM₁ and PM₄ as influential predictors of health-related outcomes, underscoring the need to expand national air quality standards beyond PM2.5 and PM₁₀. Granger-causal links, residual diagnostics, and health symptom anomalies revealed significant associations between particulate pollution and respiratory, cardiovascular, and visual symptoms, particularly in Mahendragarh. Policy insights emphasize cleaner fuel adoption, improved ventilation, and awareness campaigns to mitigate risk among vulnerable, low-income households. By integrating machine learning with epidemiological modeling, this study provides robust, location-specific evidence to support targeted environmental health interventions in under-monitored regions. A key innovation of this study lies in the joint monitoring and modeling of PM₁ and PM₄ alongside conventional PM₂.₅ and PM₁₀ using explainable ML and causal inference. This framework captures nonlinear exposure–response patterns and improves predictive accuracy while providing mechanistic insight into particle-size-specific health risks. The results offer actionable evidence for clean fuel transition, household ventilation improvements, and community-level air quality management in semi-urban and rural settings. Integration of high-resolution particulate monitoring, machine learning, and causal inference reveals strong links between PM₁-PM₁₀ exposure and cardiopulmonary and ocular symptoms in semi-urban India, highlighting PM₁ and PM₄ as key predictors for targeted interventions.
- Research Article
6
- 10.1038/s41598-025-11116-5
- Jul 15, 2025
- Scientific reports
This research presents a novel approach to improving electric power quality using semiconductor devices by integrating Machine Learning (ML), Deep Learning (DL), and advanced control strategies. The research addresses key power quality challenges - including voltage sags, swells, harmonics, and transient disturbances - through a data-driven framework that combines traditional control techniques with adaptive learning models. A variety of algorithms, including Support Vector Machines (SVM), Random Forests, Neural Networks, Convolutional Neural Networks (CNN), and Long Short-Term Memory (LSTM) networks, were tested using real-time data. The results showed notable differences in performance, with deep learning models, especially LSTM, proving to be more accurate and dependable in identifying and forecasting power quality issues. In contrast, traditional ML models like SVM and Random Forest had difficulties with class imbalance, resulting in lower precision and recall. DL models, however, managed these challenges effectively, with CNN achieving a precision of 91.8% and LSTM attaining perfect accuracy (100%) and a recall of 94.5%. The study also highlighted the complications of handling imbalanced datasets, as indicated by classification warnings, emphasizing the importance of improved preprocessing and model adjustments for reliable predictions. The execution times varied significantly, with traditional control systems being faster but less capable in identifying complex patterns compared to the computationally intensive DL models. These findings highlight the promise of hybrid systems that integrate both traditional and data-driven control strategies to achieve adaptive and dependable power quality management. Both simulations and real-world experiments support the effectiveness of this hybrid method, suggesting a strong foundation for intelligent power quality solutions in future smart grid applications. The research concludes that although deep learning models offer superior accuracy and predictive power for complex power quality scenarios, practical deployment requires careful balancing of computational demands and addressing class distribution challenges.
- Research Article
5
- 10.37745/ejaafr.2013/vol12n45480
- Mar 15, 2024
- European Journal of Accounting, Auditing and Finance Research
The integration of machine learning (ML) algorithms in Anti-Money Laundering (AML) practices has garnered significant attention due to its potential to enhance the detection and prevention of illicit activities in the cryptocurrency ecosystem. This systematic literature review analysed the effectiveness of integrating ML algorithms in detecting and preventing crypto laundering activities, identify the most frequently used ML algorithms, examine trends in publication and research methodologies, and discuss key challenges and constraints associated with integrating ML technologies into AML frameworks. A comprehensive search strategy was employed to identify relevant studies, resulting in the inclusion of 52 articles published between 2019 and 2023. The findings reveal a growing interest in the field, with a notable increase in publications in recent years. Traditional ML models such as Logistic Regression, Random Forest, and Support Vector Machine (SVM) remain prevalent, while deep learning models like Multilayer Perceptrons (MLP) and Long Short-Term Memory (LSTM) networks are gaining popularity. Graph Convolutional Networks (GCNs) have emerged as a significant area of exploration, particularly in the context of graph data analysis in cryptocurrencies. Despite advancements in ML, cryptocurrencies continue to pose a high risk of money laundering due to the practical challenge of implementation ownership of the various ML models. Future research should focus on how these challenges will be addressed to ensure the effective and sustainable use of ML technologies in real-world AML practices.
- Research Article
47
- 10.1016/j.apenergy.2023.121052
- Apr 13, 2023
- Applied Energy
Electricity peak shaving for commercial buildings using machine learning and vehicle to building (V2B) system
- Peer Review Report
- 10.5194/acp-2023-38-rc2
- Mar 31, 2023
<strong class="journal-contentHeaderColor">Abstract.</strong> As air pollution is regarded as the single largest environmental health risk in Europe it is important that communication to the public is up-to-date, accurate and provides means to avoid exposure to high air pollution levels. Long- as well as short-term exposure to outdoor air pollution is associated with increased risks of mortality and morbidity. Up-to-date information on present and coming days’ air quality help people avoid exposure during episodes with high levels of air pollution. Air quality forecasts can be based on deterministic dispersion modelling, but to be accurate this requires detailed information on future emissions, meteorological conditions and process oriented dispersion modelling. In this paper we apply different machine learning (ML) algorithms – Random forest (RF), Extreme Gradient Boosting (XGB) and Long-Short Term Memory (LSTM) – to improve 1-, 2- and 3-day deterministic forecasts of PM<sub>10</sub>, NO<sub>x</sub>, and O<sub>3</sub> at different sites in Greater Stockholm, Sweden. It is shown that the deterministic forecasts can be significantly improved using the MLs but that the degree of improvement of the deterministic forecasts depends more on pollutant and site than on what machine learning (ML) algorithm is applied. Deterministic forecasts of PM<sub>10</sub> is improved by the MLs through the input of lagged measurements and Julian day partly reflecting seasonal variations not properly parameterised in the deterministic forecasts. A systematic discrepancy by the deterministic forecasts in the diurnal cycle of NO<sub>x</sub> is removed by the MLs considering lagged measurements and calendar data like hour of the day and weekday reflecting the influence of local traffic emissions. For O<sub>3</sub> at the urban background site the local photochemistry not properly accounted for by the relatively coarse Copernicus Atmosphere Monitoring Service ensemble model (CAMS) used here for forecasting O<sub>3</sub>, but compensated using the MLs by taking lagged measurements into account. The machine learning models performed similarly well for the sites and pollutants. Performance measures like Pearson correlation, root mean square error (RMSE), mean absolute percentage error (MAPE) and mean absolute error (MAE), typically differed less than 30 % between ML models. At the urban background site, the deviations between modelled and measured concentrations (RMSE errors) are smaller than uncertainties in the measurements estimated according to recommendations by the Forum for Air Quality Modeling (FAIRMODE) in the context of the air quality directives. At the street canyon sites modelled errors are higher, and similar to measurement uncertainties. Further work is needed to reduce deviations between model results and measurements for short periods with relatively high concentrations (peaks). Such peaks can be due to a combination of non-typical emissions and unfavourable meteorological conditions and may be difficult to forecast. We have also shown that deterministic forecasts of NO<sub>x</sub> at street canyon sites can be improved using MLs even if they are trained at other sites. For PM<sub>10</sub> this was only possible using LSTM. An important aspect to consider when choosing ML is that the decision tree based models (RF and XGB) can provide useful output on the importance of features that is not possible using neural network models like LSTM, and also that training and optimisation is more complex with LSTM, which could be important to consider when selecting ML algorithm in an operational forecast system. A random forest model is now implemented operationally in the forecasts of air pollution and health risks in Stockholm. Development of the tuning process and identification of more efficient predictors may make forecast more accurate.
- Peer Review Report
- 10.5194/acp-2023-38-ac1
- May 2, 2023
<strong class="journal-contentHeaderColor">Abstract.</strong> As air pollution is regarded as the single largest environmental health risk in Europe it is important that communication to the public is up-to-date, accurate and provides means to avoid exposure to high air pollution levels. Long- as well as short-term exposure to outdoor air pollution is associated with increased risks of mortality and morbidity. Up-to-date information on present and coming days’ air quality help people avoid exposure during episodes with high levels of air pollution. Air quality forecasts can be based on deterministic dispersion modelling, but to be accurate this requires detailed information on future emissions, meteorological conditions and process oriented dispersion modelling. In this paper we apply different machine learning (ML) algorithms – Random forest (RF), Extreme Gradient Boosting (XGB) and Long-Short Term Memory (LSTM) – to improve 1-, 2- and 3-day deterministic forecasts of PM<sub>10</sub>, NO<sub>x</sub>, and O<sub>3</sub> at different sites in Greater Stockholm, Sweden. It is shown that the deterministic forecasts can be significantly improved using the MLs but that the degree of improvement of the deterministic forecasts depends more on pollutant and site than on what machine learning (ML) algorithm is applied. Deterministic forecasts of PM<sub>10</sub> is improved by the MLs through the input of lagged measurements and Julian day partly reflecting seasonal variations not properly parameterised in the deterministic forecasts. A systematic discrepancy by the deterministic forecasts in the diurnal cycle of NO<sub>x</sub> is removed by the MLs considering lagged measurements and calendar data like hour of the day and weekday reflecting the influence of local traffic emissions. For O<sub>3</sub> at the urban background site the local photochemistry not properly accounted for by the relatively coarse Copernicus Atmosphere Monitoring Service ensemble model (CAMS) used here for forecasting O<sub>3</sub>, but compensated using the MLs by taking lagged measurements into account. The machine learning models performed similarly well for the sites and pollutants. Performance measures like Pearson correlation, root mean square error (RMSE), mean absolute percentage error (MAPE) and mean absolute error (MAE), typically differed less than 30 % between ML models. At the urban background site, the deviations between modelled and measured concentrations (RMSE errors) are smaller than uncertainties in the measurements estimated according to recommendations by the Forum for Air Quality Modeling (FAIRMODE) in the context of the air quality directives. At the street canyon sites modelled errors are higher, and similar to measurement uncertainties. Further work is needed to reduce deviations between model results and measurements for short periods with relatively high concentrations (peaks). Such peaks can be due to a combination of non-typical emissions and unfavourable meteorological conditions and may be difficult to forecast. We have also shown that deterministic forecasts of NO<sub>x</sub> at street canyon sites can be improved using MLs even if they are trained at other sites. For PM<sub>10</sub> this was only possible using LSTM. An important aspect to consider when choosing ML is that the decision tree based models (RF and XGB) can provide useful output on the importance of features that is not possible using neural network models like LSTM, and also that training and optimisation is more complex with LSTM, which could be important to consider when selecting ML algorithm in an operational forecast system. A random forest model is now implemented operationally in the forecasts of air pollution and health risks in Stockholm. Development of the tuning process and identification of more efficient predictors may make forecast more accurate.
- Peer Review Report
- 10.5194/acp-2023-38-rc1
- Mar 21, 2023
<strong class="journal-contentHeaderColor">Abstract.</strong> As air pollution is regarded as the single largest environmental health risk in Europe it is important that communication to the public is up-to-date, accurate and provides means to avoid exposure to high air pollution levels. Long- as well as short-term exposure to outdoor air pollution is associated with increased risks of mortality and morbidity. Up-to-date information on present and coming days’ air quality help people avoid exposure during episodes with high levels of air pollution. Air quality forecasts can be based on deterministic dispersion modelling, but to be accurate this requires detailed information on future emissions, meteorological conditions and process oriented dispersion modelling. In this paper we apply different machine learning (ML) algorithms – Random forest (RF), Extreme Gradient Boosting (XGB) and Long-Short Term Memory (LSTM) – to improve 1-, 2- and 3-day deterministic forecasts of PM<sub>10</sub>, NO<sub>x</sub>, and O<sub>3</sub> at different sites in Greater Stockholm, Sweden. It is shown that the deterministic forecasts can be significantly improved using the MLs but that the degree of improvement of the deterministic forecasts depends more on pollutant and site than on what machine learning (ML) algorithm is applied. Deterministic forecasts of PM<sub>10</sub> is improved by the MLs through the input of lagged measurements and Julian day partly reflecting seasonal variations not properly parameterised in the deterministic forecasts. A systematic discrepancy by the deterministic forecasts in the diurnal cycle of NO<sub>x</sub> is removed by the MLs considering lagged measurements and calendar data like hour of the day and weekday reflecting the influence of local traffic emissions. For O<sub>3</sub> at the urban background site the local photochemistry not properly accounted for by the relatively coarse Copernicus Atmosphere Monitoring Service ensemble model (CAMS) used here for forecasting O<sub>3</sub>, but compensated using the MLs by taking lagged measurements into account. The machine learning models performed similarly well for the sites and pollutants. Performance measures like Pearson correlation, root mean square error (RMSE), mean absolute percentage error (MAPE) and mean absolute error (MAE), typically differed less than 30 % between ML models. At the urban background site, the deviations between modelled and measured concentrations (RMSE errors) are smaller than uncertainties in the measurements estimated according to recommendations by the Forum for Air Quality Modeling (FAIRMODE) in the context of the air quality directives. At the street canyon sites modelled errors are higher, and similar to measurement uncertainties. Further work is needed to reduce deviations between model results and measurements for short periods with relatively high concentrations (peaks). Such peaks can be due to a combination of non-typical emissions and unfavourable meteorological conditions and may be difficult to forecast. We have also shown that deterministic forecasts of NO<sub>x</sub> at street canyon sites can be improved using MLs even if they are trained at other sites. For PM<sub>10</sub> this was only possible using LSTM. An important aspect to consider when choosing ML is that the decision tree based models (RF and XGB) can provide useful output on the importance of features that is not possible using neural network models like LSTM, and also that training and optimisation is more complex with LSTM, which could be important to consider when selecting ML algorithm in an operational forecast system. A random forest model is now implemented operationally in the forecasts of air pollution and health risks in Stockholm. Development of the tuning process and identification of more efficient predictors may make forecast more accurate.
- Dissertation
- 10.12794/metadc2443360
- May 1, 2025
The Dallas-Fort Worth (DFW) metroplex is one of the fastest-growing metropolitan regions in the United States and serves as the largest economic hub in the Southern United States. Despite extensive regulatory efforts, the region is classified as an ozone non-attainment area based on the National Ambient Air Quality Standards (NAAQS), posing significant public health risks due to prolonged exposure to elevated ozone levels. Ozone, a secondary pollutant, is formed through complex photochemical reactions involving precursors such as volatile organic compounds (VOCs) and nitrogen oxides (NOx) rather than directly emitted from sources. The non-linear interaction of ozone precursors in the atmosphere presents substantial challenges in developing effective ozone reduction strategies. This dissertation analyzes air pollution measurements from 2000 to 2023, collected from multiple air quality monitoring stations across the DFW metroplex, using data mining, statistical analysis, source apportionment techniques, and machine learning (ML). By leveraging advanced techniques, this study aims to enhance the understanding of spatiotemporal pollution trends and improve air quality management within the region. Since 2000, concentrations of oxides of nitrogen (NOx) and carbon monoxide (CO) – key pollutants primarily emitted from traffic and other combustion-related sources – have shown a significant decline across the DFW metroplex. However, despite this reduction in conventional urban emissions, ozone concentrations at Denton Airport South (DEN), an exurban site, Fort Worth Northwest (FWNW), a semiurban site, and Dallas Hinton (DAL), a highly urbanized site in the DFW, have exhibited only a minor reduction, yet none of the three sites consistently met the NAAQS for ozone attainment. DAL intermittently achieved attainment status, but DEN and FWNW remained in non-attainment throughout the study period. A major contributing factor to this persistent ozone issue is the Barnett Shale, a large shale gas formation adjacent to DFW, which has been a significant source of unconventional total non-methane hydrocarbons (NMHC) emissions. The mean NMHC concentration at DEN (207.33 ± 317.23 ppb-C) – located within an active shale gas region (SGR) – was found to be more than twice the levels observed at DAL (60.54 ± 49.71 ppb-C) and FWNW (80.95 ± 65.37 ppb-C). These findings indicate that emissions from shale gas activities have contributed substantially to the atmospheric VOC levels in the region, potentially contributing to elevated ozone levels despite reductions in NOx and CO emissions from urban sources. The ozone formation potential (OFP) of NMHC at DEN was overwhelmingly dominated by slow-reacting alkanes primarily emitted from natural gas sources. In contrast, DAL was impacted by alkanes, alkenes, and aromatics from conventional urban sources, such as traffic emissions. While FWNW was impacted by NMHC from a mix of urban and natural gas sources. Using the Dispersion Normalized Positive Matrix Factorization (DN-PMF) technique, a source apportionment analysis of NMHC concentrations at DEN identified eight distinct source factors, with oil and gas activities accounting for over 94% of the total measured NMHC concentrations. The top contributing sources included natural gas extraction (71.5%), a mixed source consisting of natural gas and aviation fuel combustion (8.3%), condensate production (7.7%), and crude oil extraction (6.5%). At FWNW, NMHC concentrations were influenced by seven source factors, with natural gas extraction (37.8%) and traffic emissions (21.1%) as the dominant contributors. In DAL, NMHC levels were primarily driven by traffic-related emissions, where diesel and gasoline sources combined accounted for 32.1%, while natural gas sources contributed 31%. These findings underscore the significant role of natural gas extraction and production activities affecting the measured VOC concentrations at DEN, whereas urban traffic emissions played a more prominent role along with natural gas in NMHC profiles at DAL. Meanwhile, FWNW was impacted by a mixed composition of natural gas and urban emission sources. This analysis underscores the spatiotemporal heterogeneity of air emissions over the urban and exurban areas of North Texas. Given the variations in air emissions over any region, coupled with non-linear interactions of air pollutants in the atmosphere, real-time or near-real-time forecasting of air pollution levels remains quite challenging. Air pollutant concentration forecasting models were developed using machine learning (ML) algorithms, including Artificial Neural Networks (ANN), Random Forest (RF), and Extreme Gradient Boosting (XGBoost). RF and XGBoost models were specifically applied for VOC concentration prediction at the Denton Airport South (DEN) site, utilizing oil and natural gas (ONG) production activity data. The RF model demonstrated superior performance, achieving an R² > 0.89, indicating strong predictive accuracy for VOC trends. For 8-hour ozone concentration forecasting, a novel hybrid Recurrent Neural Network (RNN) was developed using Long Short-Term Memory (LSTM) and Gated Recurrent Unit (GRU) architectures. The model was initially trained on data from FWNW, where it exhibited excellent predictive accuracy with R² > 0.97. When applied to DAL and DEN for the years 2021 through 2023, the model maintained high performance, achieving R² > 0.94 for DAL and R² > 0.91 for DEN. Future improvements to these models could involve the integration of additional domain-specific variables, such as more detailed emission inventories, to further enhance predictive capabilities. Additionally, the development of an advanced machine learning model incorporating a comprehensive source-chemical fingerprint database would facilitate automated and highly accurate source attribution, reducing the complexity associated with manual identification in source apportionment studies. Training ML models on such extensive datasets would improve source characterization and contribute to more effective air quality management strategies.
- Research Article
42
- 10.3390/en13205429
- Oct 17, 2020
- Energies
Coordinated charging of electric vehicles (EVs) improves the overall efficiency of the power grid as it avoids distribution system overloads, increases power quality, and decreases voltage fluctuations. Moreover, the coordinated charging supports flattening the load profile. Therefore, an effective coordination technique is crucial for the protection of the distribution grid and its components. The substantial power used through charging EVs has undeniable negative impacts on the power grid. Additionally, with the increasing use of EVs, an effective solution for the coordination of EVs charging, particularly when considering the anticipated proliferation of EV fast chargers, is imminently required. In this paper, different machine learning (ML) approaches are compared for the coordination of EVs charging. The ML models can predict the power to be used in EVs charging stations (EVCS). Due to its ability to use historical data to learn and identify patterns for making future decisions with minimal user intervention, ML has been utilized. ML models used in this paper are (1) Decision Tree (DT), (2) Random Forest (RF), (3) Support Vector Machine (SVM), (4) Naïve Bayes (NB), (5) K-Nearest Neighbors (KNN), (6) Deep Neural Networks (DNN), and (7) Long Short-Term Memory (LSTM). These approaches are chosen as they are classifiers known to have the leading results for multiclass classification problems. The results found shed insight on the importance of the techniques used and their high potential in providing a reliable solution for the coordinated charging of EVs, thus improving the performance of the power grid, and reducing power losses and voltage fluctuations. The use of ML provides a less complex method to coordinate EVs, in comparison with conventional optimization techniques such as quadratic programming, and the use of ML is faster as it requires less computational power. LSTM provided the best results with an accuracy of 95% for predicting the most appropriate power rating (PR) for EVCS, followed by RF, DT, DNN, SVM, KNN, and NB. Additionally, LSTM was also the model with the smallest error rate, at a value of ±0.7%, followed by RF, DT, KNN, SVM, DNN, and NB. The results obtained from the LSTM model were similar to the results obtained from past literature using quadratic programming, with the increased speed and simplicity of ML.
- Research Article
43
- 10.1007/s11269-025-04093-x
- Jan 14, 2025
- Water Resources Management
Among natural hazards, floods pose the greatest threat to lives and livelihoods. To reduce flood impacts, short-term flood forecasting can contribute to early warnings that provide communities with time to react. This manuscript explores how machine learning (ML) can support short-term flood forecasting. Using two methods [strengths, weaknesses, opportunities, and threats (SWOT) and comparative performance analysis] for different forecast lead times (1–6, 6–12, 12–24, and 24–48 h), we evaluate the performance of machine learning models in 94 journal papers from 2001 to 2023. SWOT reveals that the best short-term flood forecasting was produced by hybrid, random forest (RF), long short-term memory (LSTM), artificial neural network (ANN), and adaptive neuro-fuzzy inference system (ANFIS) approaches. The comparative performance analysis, meanwhile, favors convolutional neural network, ANFIS, multilayer perceptron, k-nearest neighbors algorithm (KNN), hybrid, LSTM, ANN, and support vector machine (SVM) at 1–6 h; hybrid, ANFIS, ANN, and LSTM at 6–12 h; SVM, hybrid, and RF at 12–24 h; and hybrid and RF at 24–48 h. In general, hybrid approaches consistently perform well across all lead times. Trends such as hybridization, model selection, input data selection, and decomposition seem to improve the accuracy of models. Furthermore, effective stand-alone ML models such as ANN, SVM, RF, genetic algorithm, KNN, and LSTM, provide better outcomes through hybridization with other ML models. By including different machine learning models and parameters such as environmental, socio-economical, and climatic parameters, the hybrid system can produce more accurate flood forecasting, making it more effective for early warning operational purposes.
- Research Article
- 10.1145/3695251
- Nov 23, 2024
- ACM Transactions on Asian and Low-Resource Language Information Processing
As the number of social networking sites grows, so do cyber dangers. Cyberbullying is harmful behavior that uses technology to intimidate, harass, or harm someone, often on social media platforms like 𝕏 (formerly known as Twitter). Machine learning is the optimal approach for cyberbullying detection on 𝕏 to process large amounts of data, identify patterns of offensive behavior, and automate the detection process for corpus of tweets. To identify cyber threats using a trained model, the boosted ensemble (BE) technique is assessed with various machine learning algorithms such as the convolutional neural network (CNN), long short-term memory (LSTM), naive Bayes (NB), decision tree (DT), support vector machine (SVM), bidirectional LSTM (BILSTM), recurrent neural network LSTM (RNN-LSTM), multi-modal cyberbullying detection (MMCD), and random forest (RF). These classifiers are trained on the vectorized data to classify the tweets to identify cyberbullying threats. The proposed framework can detect cyberbullying cases precisely on tweets. The significance of the work lies in detecting and mitigating cyber threats in real time, and it impacts in enhancing the safety and well-being of social media users by reducing instances of cyberbullying and other cyber threats. The comparative analysis is done using metrics like accuracy, precision, recall, and F1-score, and the comparison results show that the BE technique outperforms other compared algorithms with its overall performance. Respectively, the accuracy rates of CNN, LSTM, NB, DT, SVM, RF, BILSTM, and BE are 92.5%, 93.5%, 84.6%, 88%, 89.3%, 92%, 93.75%, and 96%; precision rates of CNN, LSTM, NB, DT, SVM, RF, RNN-LSTM, and BE are 90.2%, 91.3%, 88%, 85%, 86%, 91.6%, 92.1%, and 94%; recall rates of CNN, LSTM, NB, DT, SVM, RF, BILSTM, and BE are 89.8%, 90.7%, 90%, 82%, 88.67%, 89%, 91.04%, and 93.7%; and F1-scores of CNN, LSTM, NB, DT, SVM, RF, MMCD, and BE are 90.6%, 91.8%, 85%, 84.56% 87.2%, 90%, 84.6%, and 94.89%.
- Research Article
27
- 10.1109/tte.2022.3187870
- Dec 1, 2022
- IEEE Transactions on Transportation Electrification
An effective fault locating method is necessary to ensure the stable and efficient operation of solid oxide fuel cells (SOFCs). There is still a lack of a common fault locating method for locating multiple faults in SOFC systems. Therefore, this article proposes a multifault spatiotemporal locating method combining long short-term memory (LSTM) artificial neural network and causal inference. This method does not rely on the SOFC mechanism model and does not require a large amount of fault data. This method has good migratory characteristics and can be used with different systems. This method first reconstructs the experimental data by LSTM and locates the fault occurrence time according to the reconstruction error. Then, the space where the fault occurred is located by the causal inference method. At the same time, multiple locating methods are compared. Finally, a performance optimization method is adopted from the system level to improve the efficiency of the system. From the comparison results, it can be seen that the scheme proposed in this article is able to locate different faults in time and space with an accuracy of 92.6%. In addition, the system efficiency can be improved by 18.7% after the corresponding optimization methods are adopted.
- Research Article
45
- 10.1109/access.2024.3404651
- Jan 1, 2024
- IEEE Access
Advancements in the agricultural sector are essential because the need for food is rising as the population of the world is expanding day by day. Traditional agricultural practices are not able to fulfill these needs. Furthermore, these practices are manual and are not optimized resulting in the wastage of resources, which is not suitable for resource-constrained agricultural environments. Besides this, the Internet of Things (IoT) network is playing an important role in the modern farming system. In this paper, we introduce an innovative IoT-enabled hybrid model for smart agriculture, integrating Machine Learning (ML) and Artificial Intelligence (AI) algorithms to provide a cost-effective and reliable decision-making system. Furthermore, we introduce a robust anomaly detection mechanism while applying the capabilities of Multilayer Perceptron (MLP), Naïve Bayes, and Support Vector Machine (SVM) on the dry beans’ dataset. Hybrid models, combining neural networks with Random Forest and SVM, were also explored for anomaly detection in the dataset. Furthermore, deep learning models known as MobileNetV2, VGG16, and InceptionV3 are used for the classification of soil type datasets. The hybrid deep learning models were also developed, incorporating InceptionV3 with Long Short-Term Memory (LSTM) and VGG16 with fully connected dense layers. Two types of data sets are used in this study, which are the dry beans dataset (2021) and soil type Dataset (2024). Both datasets contain images. The ML techniques are applied to these datasets for anomaly detection. The simulations results show that the classification performance of the MobileNetV2 model, it has an accuracy and recall of 0.97. It shows that the model can correctly identify the soil type around 97%. On the other hand, the hybrid model combining random forest and neural network achieved an accuracy of 92%, further validating the effectiveness of our approach. Furthermore, the SVM model achieves an impressive overall accuracy of 0.93. Additionally, this accuracy is further enhanced with the integration of SVM and neural networks. Similarly, the hybrid model combining inception V3 with the LSTM layer exhibits a notable accuracy of 0.91, highlighting its efficiency in accurately classifying various instances. Lastly, the hybrid model employing random forest and neural network architecture achieves a commendable accuracy of 92%.
- Research Article
3
- 10.3390/automation5030021
- Aug 1, 2024
- Automation
This study addresses the integration of machine learning (ML) with supervisory control and data acquisition (SCADA) systems to enhance predictive maintenance and operational efficiency in oil well monitoring. We investigated the applicability of advanced ML models, including Long Short-Term Memory (LSTM), Bidirectional LSTM (BiLSTM), and Momentum LSTM (MLSTM), on a dataset of 21,644 operational records. These models were trained to predict a critical operational parameter, FlowRate, which is essential for operational integrity and efficiency. Our results demonstrate substantial improvements in predictive accuracy: the LSTM model achieved an R2 score of 0.9720, the BiLSTM model reached 0.9725, and the MLSTM model topped at 0.9726, all with exceptionally low Mean Absolute Errors (MAEs) around 0.0090 for LSTM and 0.0089 for BiLSTM and MLSTM. These high R2 values indicate that our models can explain over 97% of the variance in the dataset, reflecting significant predictive accuracy. Such performance underscores the potential of integrating ML with SCADA systems for real-time applications in the oil and gas industry. This study quantifies ML’s integration benefits and sets the stage for further advancements in autonomous well-monitoring systems.
- Research Article
9
- 10.13031/ja.15995
- Jan 1, 2024
- Journal of the ASABE
Highlights A 10-year dataset of time-series MODIS imagery and in situ Chl- a concentration were curated for Lake Okeechobee. LSTM significantly outperformed KNN, SVR, and RF for Chl- a prediction and subsequent HAB detection. The optimal window length was found to be 13 days with a 4-day temporal resolution for the LSTM model. KNN, SVR, and RF models were not effective at utilizing the temporal dynamics of the input features. Abstract. Harmful algal blooms (HABs) in inland waterbodies are a global concern due to their negative impact on human, animal, and ecosystem health. Chlorophyll-a (Chl-a) concentration is an important water quality parameter for monitoring HABs. While statistical and machine learning (ML) models have been widely studied to predict Chl-a concentration and HABs based on single-time-point satellite data, this work assessed whether long short-term memory (LSTM) can improve both tasks by leveraging temporal features in time-series MODIS satellite images compared to three classical ML models, including k-nearest neighbor (KNN), support vector regression (SVR), and random forest (RF). A dataset of daily MODIS images and monthly in situ Chl-a concentration measurements from 2011 to 2020 was curated for Lake Okeechobee, Florida. A window size of 13 days with a temporal resolution of four days was found to produce the optimal performance for LSTM, which significantly outperformed KNN, SVR, and RF for Chl-a prediction with a root mean square error of 11.95 µg/L, a mean absolute error of 8.55 µg/L, and a R 2 value of 0.43. The superior performance of LSTM for Chl-a prediction was likely due to its ability to leverage the temporal dynamics in the features associated with HAB development. The Chl-a predictions were further used to determine HAB events, showing better accuracy and a significantly higher F1 score for LSTM over the other models. The study suggested that combining LSTM with high-temporal-resolution time-series data should be preferred over applying common ML models on time-series or single-time-point remote sensing data for Chl-a and HAB monitoring. Keywords: Cyanobacteria, LSTM, Machine Learning, Remote Sensing, Water Quality.
- Research Article
- 10.11591/ijai.v14.i6.pp4828-4837
- Dec 1, 2025
- IAES International Journal of Artificial Intelligence (IJ-AI)
Elevated temperatures in urban areas relative to surrounding rural areas, known as the urban heat island (UHI) effect, constitute a pressing challenge to urban sustainability, public health, and energy efficiency. With a comprehensive global dataset from NASA's Socioeconomic Data and Applications Center (SEDAC) that encompasses land surface temperature (LST) and different urban characteristics, this study investigates the UHI phenomenon. The UHI intensity was predicted using advanced machine learning models, random forest, extreme gradient boosting (XGBoost), light gradient boosting machine (LightGBM), multilayer perceptron (MLP), and long short-term memory (LSTM) with attention mechanism. The LSTM with attention achieved top R2:0.9998 (day) and 0.9992 (night). Key landscape metrics include urban area size, population, and location. We analyzed spatial temporal UHI patterns to identify local factors like geometry and vegetation. These findings are critical for urban planners and policy makers to identify targeted mitigation options, including green space expansion, the use of low thermal mass, and urban climate resilience strategies. These results advance predictive modeling, supporting resilient, and sustainable cities.