Accelerate Literature Icon
Want to do a literature review? Try our new Literature Review workflow

Gujarati Task Oriented Dialogue Slot Tagging Using Deep Neural Network Models

  • TL;DR
  • Abstract
  • Literature Map
  • Similar Papers
TL;DR

This study evaluates deep neural network models for Gujarati dialogue slot tagging, finding that BiLSTM-based models outperform LSTM variants by approximately 2% in F1-measure, with BiLSTM-CRF achieving the highest performance, enhancing spoken language understanding in Gujarati.

Abstract
Translate article icon Translate Article Star icon

In this paper, the primary focus is of Slot Tagging of Gujarat Dialogue, which enables the Gujarati language communication between human and machine, allowing machines to perform given task and provide desired output. The accuracy of tagging entirely depends on bifurcation of slots and word embedding. It is also very challenging for a researcher to do proper slot tagging as dialogue and speech differs from human to human, which makes the slot tagging methodology more complex. Various deep learning models are available for slot tagging for the researchers, however, in the instant paper it mainly focuses on Long Short-Term Memory (LSTM), Convolutional Neural Network - Long Short-Term Memory (CNN-LSTM) and Long Short-Term Memory – Conditional Random Field (LSTM-CRF), Bidirectional Long Short-Term Memory (BiLSTM), Convolutional Neural Network - Bidirectional Long Short-Term Memory (CNN-BiLSTM) and Bidirectional Long Short-Term Memory – Conditional Random Field (BiLSTM-CRF). While comparing the above models with each other, it is observed that BiLSTM models performs better than LSTM models by a variation ~2% of its F1-measure, as it contains an additional layer which formulates the word string to traverse from backward to forward. Within BiLSTM models, BiLSTM-CRF has outperformed other two Bi-LSTM models. Its F1-measure is better than CNN-BiLSTM by 1.2% and BiLSTM by 2.4%.KeywordsSpoken Language Understanding (SLU)Long Short-Term Memory (LSTM)Slot taggingBidirectional Long Short-Term Memory (BiLSTM)Convolutional Neural Network - Bidirectional Long Short-Term Memory (CNN-BiLSTM)Bidirectional Long Short-Term Memory (BiLSTM-CRF)

Similar Papers
  • Research Article
  • Cite Count Icon 6
  • 10.1080/15567036.2021.1925379
Application of stacked and bidirectional long short-term memory deep learning models for wind speed forecasting at an offshore site
  • Aug 26, 2021
  • Energy Sources, Part A: Recovery, Utilization, and Environmental Effects
  • Bharat Kumar Saxena + 2 more

Very short-term offshore wind speed forecasting by application of Stacked long short-term memory (LSTM) and Bidirectional LSTM deep learning models is done in this work. Wind speed data of two different offshore sites located in two different continents are used for testing the models. Performance is measured on the basis of accuracy of forecasting and computational time. The effectiveness of Stacked LSTM and Bidirectional LSTM models is also validated by comparing their performance with convolutional neural network, convolutional neural network-long short-term memory network, multi-layer perceptron, and rolling forecasting auto regressive integrated moving average models. Results of forecasting error confirm that Stacked LSTM model is better than other compared models in forecasting very short-term offshore wind speed. Mean absolute percentage error (MAPE) of wind speed forecasting by Stacked LSTM model is 4.59% at Anholt (Denmark) and 3.62% at Dhanushkodi (India) sites. From comparison of MAPE of Stacked LSTM model with that of eight other latest existing models in literature, it can be concluded that Stacked LSTM model is superior to many other existing models.

  • Research Article
  • 10.52783/jisem.v10i15s.2475
Evaluating the Effectiveness of CNN, LSTM, and Bi-LSTM Models in Classifying Twitter Sentiments
  • Mar 4, 2025
  • Journal of Information Systems Engineering and Management
  • Naveen P

Introduction: Twitter sentiment analysis is an essential tool for understanding public opinions and extracting valuable insights from social media discussions. However, the informal, concise nature of tweets, along with their context-dependent language, poses significant challenges for accurate sentiment classification. Deep learning techniques have shown promise in addressing these challenges due to their ability to learn complex patterns and contextual relationships in data. This study explores the application of Convolutional Neural Networks (1D-CNN), Long Short-Term Memory (LSTM), and Bidirectional LSTM (Bi-LSTM) models for classifying tweet sentiments into four categories: positive, negative, neutral, and irrelevant. Objectives: The primary objective of this study is to evaluate the effectiveness of deep learning models, including 1D-CNN, LSTM, and Bi-LSTM, in classifying tweet sentiments into categories such as positive, negative, neutral, and irrelevant. By analyzing the performance of these models, the study aims to compare their accuracy, precision, recall, and F1-score to determine their strengths and limitations. Additionally, the research seeks to identify the most suitable model capable of effectively capturing the sequential and contextual nature of tweets, addressing the unique challenges posed by the informal and context-dependent language of social media data. Methods: The study employs a publicly available Twitter sentiment dataset comprising 73,906 tweets related to general Twitter discussions. The dataset undergoes pre-processing steps, including tokenization, stopword removal, and handling emoticons and hashtags, to ensure clean input for training the models. Tweets are represented using pre-trained word embeddings, specifically GloVe and Word2Vec, which provide rich semantic information by capturing word meanings in context. Individual models—1D-CNN, LSTM, and Bi-LSTM—are trained and evaluated using this pre-processed data. Performance metrics, including accuracy, precision, recall, and F1-score, are calculated to compare the models' effectiveness. Results: The experimental analysis demonstrates that each model has unique strengths in sentiment classification. However, the 1D-CNN outperforms both LSTM and Bi-LSTM models, achieving superior results in capturing both the sequential and contextual information inherent in tweet data. This highlights its efficiency and suitability for Twitter sentiment analysis. Conclusions: This study underscores the potential of deep learning techniques for sentiment analysis in social media. Among the evaluated models, 1D-CNN proves to be the most effective, offering a robust approach to handling the complexities of tweet sentiment classification. Future work could explore hybrid models and further optimization techniques to enhance sentiment analysis performance.

  • Conference Article
  • 10.1109/icaaic56838.2023.10140404
Neural Network-based Approach to Predict Protein Secondary Structure
  • May 4, 2023
  • Arifur Rahman + 2 more

Protein Secondary structure prediction is an emerging topic in bioinformatics to understand briefly the functions of protein and their role in drug invention, medicine and biology. In our research we have applied two recurrent neural network based approach Bi-LSTM (Bidirectional Long Short-Term Memory) and LSTM (Long Short-Term Memory). Our research was focused on primary structure up to 134 in length of amino acids. Initially our proposed model produced a ‘Indexed Lexicon of corpus’ using tri-gram conversion for primary structure strings. Each primary structure tri-gram transformed snippets is substituted with its associated index mentioned in ‘Indexed corpus’. The indexed parameter vector inputted into our proposed Bi-LSTM and LSTM model. We got best accuracy when we have used two Bi-LSTM and three LSTM layers respectively in Bi-LSTM and LSTM models. To prevent biasness and minimize overfitting problem we have utilized two dropout layers for each of Bi-LSTM and LSTM model. We have operated our model on ccPDB 2.0 benchmark dataset. There is total eight states protein secondary structure in this dataset. For this sst8 secondary structure we have achieved 83.24% accuracy for our proposed LSTM model and 89.10% accuracy for our Bi-LSTM model. We have configured our model to run for 50 epochs with batch size 64. For compilation of our models we have utilized ‘adam’ optimizer and the ‘categorical crossentropy’ loss function. To make dataset balanced to our model we have also employed 5-fold cross validation.

  • Research Article
  • Cite Count Icon 7
  • 10.26555/ijain.v10i1.1170
Emergency sign language recognition from variant of convolutional neural network (CNN) and long short term memory (LSTM) models
  • Feb 29, 2024
  • International Journal of Advances in Intelligent Informatics
  • Muhammad Amir As'Ari + 2 more

Sign language is the primary communication tool used by the deaf community and people with speaking difficulties, especially during emergencies. Numerous deep learning models have been proposed to solve the sign language recognition problem. Recently. Bidirectional LSTM (BLSTM) has been proposed and used in replacement of Long Short-Term Memory (LSTM) as it may improve learning long-team dependencies as well as increase the accuracy of the model. However, there needs to be more comparison for the performance of LSTM and BLSTM in LRCN model architecture in sign language interpretation applications. Therefore, this study focused on the dense analysis of the LRCN model, including 1) training the CNN from scratch and 2) modeling with pre-trained CNN, VGG-19, and ResNet50. Other than that, the ConvLSTM model, a special variant of LSTM designed for video input, has also been modeled and compared with the LRCN in representing emergency sign language recognition. Within LRCN variants, the performance of a small CNN network was compared with pre-trained VGG-19 and ResNet50V2. A dataset of emergency Indian Sign Language with eight classes is used to train the models. The model with the best performance is the VGG-19 + LSTM model, with a testing accuracy of 96.39%. Small LRCN networks, which are 5 CNN subunits + LSTM and 4 CNN subunits + BLSTM, have 95.18% testing accuracy. This performance is on par with our best-proposed model, VGG + LSTM. By incorporating bidirectional LSTM (BLSTM) into deep learning models, the ability to understand long-term dependencies can be improved. This can enhance accuracy in reading sign language, leading to more effective communication during emergencies.

  • PDF Download Icon
  • Research Article
  • Cite Count Icon 189
  • 10.3390/rs12162655
Rice Crop Detection Using LSTM, Bi-LSTM, and Machine Learning Models from Sentinel-1 Time Series
  • Aug 18, 2020
  • Remote Sensing
  • Hugo Crisóstomo De Castro Filho + 9 more

The Synthetic Aperture Radar (SAR) time series allows describing the rice phenological cycle by the backscattering time signature. Therefore, the advent of the Copernicus Sentinel-1 program expands studies of radar data (C-band) for rice monitoring at regional scales, due to the high temporal resolution and free data distribution. Recurrent Neural Network (RNN) model has reached state-of-the-art in the pattern recognition of time-sequenced data, obtaining a significant advantage at crop classification on the remote sensing images. One of the most used approaches in the RNN model is the Long Short-Term Memory (LSTM) model and its improvements, such as Bidirectional LSTM (Bi-LSTM). Bi-LSTM models are more effective as their output depends on the previous and the next segment, in contrast to the unidirectional LSTM models. The present research aims to map rice crops from Sentinel-1 time series (band C) using LSTM and Bi-LSTM models in West Rio Grande do Sul (Brazil). We compared the results with traditional Machine Learning techniques: Support Vector Machines (SVM), Random Forest (RF), k-Nearest Neighbors (k-NN), and Normal Bayes (NB). The developed methodology can be subdivided into the following steps: (a) acquisition of the Sentinel time series over two years; (b) data pre-processing and minimizing noise from 3D spatial-temporal filters and smoothing with Savitzky-Golay filter; (c) time series classification procedures; (d) accuracy analysis and comparison among the methods. The results show high overall accuracy and Kappa (>97% for all methods and metrics). Bi-LSTM was the best model, presenting statistical differences in the McNemar test with a significance of 0.05. However, LSTM and Traditional Machine Learning models also achieved high accuracy values. The study establishes an adequate methodology for mapping the rice crops in West Rio Grande do Sul.

  • Conference Article
  • Cite Count Icon 1486
  • 10.1109/bigdata47090.2019.9005997
The Performance of LSTM and BiLSTM in Forecasting Time Series
  • Dec 1, 2019
  • Sima Siami-Namini + 2 more

Machine and deep learning-based algorithms are the emerging approaches in addressing prediction problems in time series. These techniques have been shown to produce more accurate results than conventional regression-based modeling. It has been reported that artificial Recurrent Neural Networks (RNN) with memory, such as Long Short-Term Memory (LSTM), are superior compared to Autoregressive Integrated Moving Average (ARIMA) with a large margin. The LSTM-based models incorporate additional “gates” for the purpose of memorizing longer sequences of input data. The major question is that whether the gates incorporated in the LSTM architecture already offers a good prediction and whether additional training of data would be necessary to further improve the prediction. Bidirectional LSTMs (BiLSTMs) enable additional training by traversing the input data twice (i.e., 1) left-to-right, and 2) right-to-left). The research question of interest is then whether BiLSTM, with additional training capability, outperforms regular unidirectional LSTM. This paper reports a behavioral analysis and comparison of BiLSTM and LSTM models. The objective is to explore to what extend additional layers of training of data would be beneficial to tune the involved parameters. The results show that additional training of data and thus BiLSTM-based modeling offers better predictions than regular LSTM-based models. More specifically, it was observed that BiLSTM models provide better predictions compared to ARIMA and LSTM models. It was also observed that BiLSTM models reach the equilibrium much slower than LSTM-based models.

  • Research Article
  • Cite Count Icon 31
  • 10.1108/aeat-05-2022-0132
A new proposal for the prediction of an aircraft engine fuel consumption: a novel CNN-BiLSTM deep neural network model
  • Mar 7, 2023
  • Aircraft Engineering and Aerospace Technology
  • Sedat Metlek

PurposeThe purpose of this study is to develop and test a new deep learning model to predict aircraft fuel consumption. For this purpose, real data obtained from different landings and take-offs were used. As a result, a new hybrid convolutional neural network (CNN)-bi-directional long short term memory (BiLSTM) model was developed as intended.Design/methodology/approachThe data used are divided into training and testing according to the k-fold 5 value. In this study, 13 different parameters were used together as input parameters. Fuel consumption was used as the output parameter. Thus, the effect of many input parameters on fuel flow was modeled simultaneously using the deep learning method in this study. In addition, the developed hybrid model was compared with the existing deep learning models long short term memory (LSTM) and BiLSTM.FindingsIn this study, when tested with LSTM, one of the existing deep learning models, values of 0.9162, 6.476, and 5.76 were obtained for R2, root mean square error (RMSE), and mean absolute percentage error (MAPE), respectively. For the BiLSTM model when tested, values of 0.9471, 5.847 and 4.62 were obtained for R2, RMSE and MAPE, respectively. In the proposed hybrid model when tested, values of 0.9743, 2.539 and 1.62 were obtained for R2, RMSE and MAPE, respectively. The results obtained according to the LSTM and BiLSTM models are much closer to the actual fuel consumption values. The error of the models used was verified against the actual fuel flow reports, and an average absolute percent error value of less than 2% was obtained.Originality/valueIn this study, a new hybrid CNN-BiLSTM model is proposed. The proposed model is trained and tested with real flight data for fuel consumption estimation. As a result of the test, it is seen that it gives much better results than the LSTM and BiLSTM methods found in the literature. For this reason, it can be used in many different engine types and applications in different fields, especially the turboprop engine used in the study. Because it can be applied to different engines than the engine type used in the study, it can be easily integrated into many simulation models.

  • Research Article
  • 10.36349/easjecs.2025.v08i02.001
Forecasting United States Dollar to Tanzania Shillings Exchange Rate Using Comparable LSTM and BiLSTM Deep Learning Models
  • Mar 1, 2025
  • East African Scholars Journal of Engineering and Computer Sciences
  • Isakwisa Gaddy Tende

Tanzania heavily depends on United States Dollar (USD) foreign currency to import various goods and services into the country. Failure to correctly forecast exchange rates between USD and Tanzanian Shillings (TZS) may pose risks such as inability to import intended goods and services, possibility of losing money in stock exchange markets and other investment businesses in case of unexpected currency appreciation or depreciation as well as poor investment decisions in foreign exchange markets. To address this, this study has developed and comparatively evaluated performances of LSTM (Long Short-Term Memory) and BiLSTM (Bidirectional LSTM) deep learning models for forecasting daily USD to TZS exchange rates. The findings reveal that, BiLSTM model outperforms LSTM in forecasting daily USD to TZS exchange rates, achieving a MAPE (Mean Absolute Percentage Error) score of 0.363 on test set (unseen data) compared to a MAPE score of 1.471 achieved by LSTM model. This study recommends to the prospective Artificial Intelligence (AI) researchers and software developers to use BiLSTM instead of LSTM model to forecast (predict) USD to TZS exchange rates. Also, this study has developed USD to TZS exchange rates dataset which can be used by AI researchers, saving them time and costs involved with creating datasets from scratch. This study has also developed ready to use BiLSTM and LSTM models which can be used by Tanzanian business men and women involved in stock exchange markets, foreign exchange markets and other businesses, to predict daily USD to TZS exchange rates and make appropriate business and investment decisions.

  • Research Article
  • Cite Count Icon 14
  • 10.28991/esj-2024-08-05-013
An Explainable Deep Learning Approach for Classifying Monkeypox Disease by Leveraging Skin Lesion Image Data
  • Oct 1, 2024
  • Emerging Science Journal
  • Andino Maseleno + 2 more

According to the World Health Organization's (WHO) external situation report on the multi-country outbreak of Monkeypox in 2023, from 11 countries in Southeast Asia Regions, Thailand recorded the highest reported cases, totaling 461. The ongoing Monkeypox outbreak has raised significant public health concerns due to its rapid spread across several nations. Early detection and diagnosis are imperative for effectively treating and controlling Monkeypox. Given this context, this study aimed to determine the most efficient model for detecting Monkeypox by employing interpretable deep learning techniques. This study utilizes deep learning techniques to diagnose Monkeypox based on images of skin lesions. We evaluate based on four models—convolutional neural network (CNN), gated recurrent unit (GRU), long short-term memory (LSTM), and bidirectional long short term memory (BiLSTM)—using a publicly available dataset. Additionally, we incorporate Local Interpretable Model-Agnostic Explanations (LIME) and techniques for explainable AI, facilitating visual interpretation of model predictions for healthcare practitioners. The CNN model's performance and LSTM model's performance have an accuracy of 100%, while the GRU model's performance and BiLSTM model's performance have an accuracy of 99.88% and 99.45%. Our findings demonstrate the effectiveness of deep learning models, including the suggested CNN model leveraging the pre-trained MobileNetV2 and LSTM. These models can play a pivotal role in combating the Monkeypox virus. Doi: 10.28991/ESJ-2024-08-05-013 Full Text: PDF

  • PDF Download Icon
  • Research Article
  • Cite Count Icon 2
  • 10.1007/s11042-025-20787-1
A novel intelligent healthcare system based on efficient deep learning models for brain stroke prediction
  • Apr 15, 2025
  • Multimedia Tools and Applications
  • Saeed Mohsen + 1 more

Smart hospitals should equip patients with wearable devices capable of monitoring and predicting brain strokes using deep learning (DL) techniques. The latest advancements in the accuracy of DL have the potential to make a substantial impact in addressing the issues associated with predicting brain strokes. However, further enhancements are required to achieve even higher levels of precision in DL methods. This research proposes an innovative intelligent healthcare system (IHS) that utilizes a real-time web application. The IHS is specifically engineered to monitor and forecast cerebral stroke occurrences in individuals. Also, this work presents three DL models: long short-term memory (LSTM), gated recurrent unit (GRU), and bidirectional long short-term memory (BiLSTM), which are used to predict brain stroke. The three models are trained using a dataset collected from 4,981 individuals. This dataset consists of two categories—normal and stroke, and it comprises eleven distinct characteristics: gender, age, heart disease (HD), hypertension, marital status (MS), type of residence (RT), average glucose level (AGL), type of work (TW), body mass index (BMI), smoking status (SS), and stroke. The DL models suggested in this study are constructed using the Keras library, employing a hyperparameter tuning technique to maximize accuracy. The effectiveness of the three models is evaluated using precision-recall (PR) curves and a normalized error matrix (NEM). The results indicate that the BiLSTM model outperforms both the GRU and LSTM models in terms of efficiency for predicting brain stroke. The BiLSTM achieves the most efficiency, with a testing accuracy (TA) of 100%, followed by the LSTM with a TA of 99.90%. However, the GRU exhibits the lowest TA of 99.80%. The testing loss (TL) rates for the LSTM, GRU, and BiLSTM models are 0.0075, 0.022, and 0.0001, respectively. Additionally, the BiLSTM model achieves sensitivity, accuracy, F1-score, and area under the PR curves of 100%. The proposed DL models with IHS can assist physicians in efficiently and precisely diagnosing persons with brain strokes, enabling them to make prompt and accurate decisions.

  • PDF Download Icon
  • Research Article
  • Cite Count Icon 2
  • 10.1088/1742-6596/1972/1/012098
Deep Learning with Bidirectional Long Short-Term Memory for traffic flow Prediction
  • Jul 1, 2021
  • Journal of Physics: Conference Series
  • Song Xue + 3 more

With the development of cities, the total number of trucks has increased year by year. Traffic flow forecasting has become an indispensable part of the cargo transportation industry and directly affects the development of the transportation industry. In the field of traffic flow prediction, Long Short-Term Memory(LSTM) model has advantages in processing time series, but it cannot extract the periodicity in time series. Therefore, in this experiment, a Bidirectional Long Short-Term Memory (BLSTM) model was constructed to predict traffic flow in the road network. It is worth mentioning that this article considers the non-parametric model autoregressive integrated moving average model (ARIMA) and the parametric model recurrent neural network (RNN) to compare and analyze with LSTM. Data from Guangwu Toll Station, Zhengzhou city, China were used to calibrate and evaluate the models. The experimental results show that the performance of RNN based on deep learning such as BLSTM and LSTM model is better than that of ARIMA. In order to better illustrate the advantages of BLSTM model, we comprehensively considered the performance effects of four models under morning peak, evening peak and flat peak. Experiments have proved that BLSTM has good nonlinear fitting ability and anti-noise ability, and the average prediction accuracy reaches 92.873%.

  • Research Article
  • Cite Count Icon 1
  • 10.3389/fbuil.2025.1699466
Estimating shield tunnel boring machine penetration rate in mixed face conditions: feature selection and multicollinearity effects on machine and deep learning models
  • Nov 19, 2025
  • Frontiers in Built Environment
  • Jitendra Khatti + 1 more

This research compares the support vector machine (SVM), gene expression programming (GEP), feedforward neural network (FFNN), gated recurrent unit (GRU), long short-term memory (LSTM), support vector regressor (SVR), and bidirectional long short-term memory (BiLSTM) models in predicting penetration (PR) rate of earth pressure balance shield tunnel boring machine (E TBM ). A dataset has been compiled using the cutterhead rotation speed (CRS), mean thrust (F/A), mean cutterhead torque (T/D 3 ), upper earth pressure (UEP), lower earth pressure (LEP), and torque penetration index (TPI) features of 1,197 E TBM events. The presence of multicollinearity was analyzed using the variance inflation factor (VIF) method. It was observed that CRS, F/A, T/D 3 , UEP, LEP, and TPI have weak, moderate, considerable, moderate, problematic, and considerable multicollinearity, respectively. The performance (R) comparison revealed that the BiLSTM models predicted PR (=1.0000 in testing and validation) with higher performance than SVM, SVR, GEP, FFNN, GRU, and LSTM models. In addition, the score analysis (=285), error characteristics curve (=7.03E-07), generalizability (m and n < 0.00), Wilcoxon test (confidence = 95.02%), uncertainty analysis (first rank), Anderson-Darling test (accept the normality hypothesis), and objective function criterion (=0.0003) presented that the BiLSTM model is an optimal performance computational model in predicting PR of E TBM . It was also noted that the CRS, F/A, T/D 3 , UEP, LEP, and TPI features are more reliable for accurately predicting PR.

  • Conference Article
  • Cite Count Icon 3
  • 10.1109/iscmi59957.2023.10458486
Classifying DNS over HTTPS Malicious/Benign Traffic Using Deep Learning Models
  • Nov 25, 2023
  • Mandar Chougule + 7 more

As we live in an era where privacy over the Internet has become rudimentary, protocols like DNS over HTTPS (DoH) and DNS over TLS (DoT), which promote encryption, have become popular. While these protocols were introduced to overcome the drawbacks of DNS protocol, even DoH has some security issues that need to be tackled to prevent any misuse. Herein, we implemented deep learning models to classify DNS over HTTPS traffic and found the most efficient method in regard to time-required complexity and computational requirements. Previous studies have used a variety of features from datasets to identify malicious activities. Although machine learning and deep learning models are commonly used, they require more human intervention. These models are also more computationally complex, as one is required to tune the model and its parameters for accurate results. In comparison, some deep learning models are more efficient as they work well without any human intervention and are capable of parameter tuning by themselves. In this work, we used the CIRA-CIC-DoHBrw-2020 dataset and performed data imbalance handling, one hot encoding, and feature selection to create a model that can be used for a more generalized environment. We implemented long short-term memory (LSTM), bidirectional LSTM (BiLSTM), and gated recurrent unit (GRU) models to classify DoH traffic with high accuracy. Although the mentioned models produced good accuracy, the BiLSTM model performs better than the LSTM model in the time taken for prediction and accuracy; the GRU model outperformed both LSTM and BiLSTM models in terms of accuracy, computation time, and computation complexity. Hence, it is more efficient than both LSTM and BiLSTM models.

  • Conference Article
  • Cite Count Icon 2
  • 10.1109/icccnt61001.2024.10725023
Exploring Authorial Style in Bangla Literature: LSTM and Bi-LSTM -Based Author Detection
  • Jun 24, 2024
  • Pronoy Kumar Mondal + 5 more

Authorship attribution, a critical task in computational etymology, includes distinguishing the author of a text based on complex subtleties. In order to address the issues with authorship attribution in Bengali literary texts, this paper compares two deep learning models: Bidirectional Long Short-Term Memory (Bi-LSTM), enhanced by Adam and RMSprop optimizers, and Long Short-Term Memory (LSTM). Utilizing a dataset comprising writings from nine prominent Bengali authors, we explored the adequacy of these models in distinguishing between different composing styles. Thorough preprocessing, feature engineering, and an experimental setup were all part of our methodology to assess the LSTM and Bi-LSTM models’ performance in terms of various metrics like accuracy, precision, recall, and F1 score. The results show that the Bi-LSTM model, optimized with Adam, achieved the highest testing accuracy of $\mathbf{92\%}$, illustrating its predominant capability to generalize from preparing information to inconspicuous information. Also, we watched striking contrasts in model performance between training and testing phases, proposing potential zones for enhancement in demonstration preparation and generalization. This study not only underscores the potential of LSTM and Bi-LSTM models in authorship attribution but also highlights the significance of optimizing model configurations to improve performance. Further research aims to improve these models, look into additional include sets, and expand the analysis to other languages and literary forms.

  • Research Article
  • Cite Count Icon 5
  • 10.57152/malcom.v4i4.1671
Spam Detection in YouTube Comments Using Deep Learning Models: A Comparative Study of MLP, CNN, LSTM, BiLSTM, GRU, and Attention Mechanisms
  • Oct 8, 2024
  • MALCOM: Indonesian Journal of Machine Learning and Computer Science
  • Gregorius Airlangga

This study explores the effectiveness of various deep learning models for detecting spam in YouTube comments. Six models were evaluated: Multilayer Perceptron (MLP), Convolutional Neural Network (CNN), Long Short-Term Memory (LSTM), Bidirectional LSTM (BiLSTM), Gated Recurrent Unit (GRU), and Attention mechanisms. The dataset consists of 1,956 real comments extracted from popular YouTube videos, representing both spam and legitimate messages. The preprocessing phase involved tokenization and padding of text sequences to prepare them for model input. Results reveal that the LSTM model achieved the highest test accuracy of 95.65%, outperforming other models by capturing sequential dependencies and context within comments. The CNN model also demonstrated high accuracy, underscoring the importance of local pattern recognition in text classification. While BiLSTM and Attention models offered comparable performance, their marginal improvement over LSTM indicates that sequential modeling plays a crucial role in this task. The GRU model, despite being computationally efficient, showed slightly lower accuracy compared to LSTM and BiLSTM. The MLP model, serving as a baseline, exhibited limited performance, emphasizing the need for advanced architectures in spam detection. These findings suggest that combining sequential modeling with local feature extraction could lead to more robust spam detection systems.

Save Icon
Up Arrow
Open/Close
Notes

Save Important notes in documents

Highlight text to save as a note, or write notes directly

You can also access these Documents in Paperpal, our AI writing tool

Powered by our AI Writing Assistant