Accelerate Literature Icon
Want to do a literature review? Try our new Literature Review workflow

Optimized Attention-Driven Bidirectional Convolutional Neural Network

  • TL;DR
  • Abstract
  • Literature Map
  • Similar Papers
TL;DR

This paper presents an optimized sentiment analysis method using tokenization via BERT and classification with an Attention-based Bidirectional CNN-RNN model trained by a novel Chimp Deer Hunting Optimization algorithm. The approach achieves high performance, with a precision of 93.5%, recall of 94.5%, and F-measure of 94%.

Abstract
Translate article icon Translate Article Star icon

This paper devises an optimization-based technique for sentiment analysis using the set of reviews. The major processes involved for the developed sentiment analysis approach are tokenization and sentiment classification. Initially, the input reviews are considered from the database and are subjected to the tokenization process. The tokenization process is performed using Bidirectional Encoder Representations from Transformer (BERT) where the input review data is partitioned into individual words, named as tokens. Finally, sentiment classification is carried out using Attention-based Bidirectional CNN-RNN Deep Model (ABCDM), which is trained by proposed Chimp Deer Hunting Optimization (CDHO) approach. Accordingly, the proposed CDHO algorithm is newly designed by incorporating Chimp Optimization Algorithm (ChOA) and Deer Hunting Optimization Algorithm (DHOA). The proposed CDHO-based ABCDM provided enhanced performance with highest precision of 93.5%, recall of 94.5% and F-measure of 94%.

Similar Papers
  • Research Article
  • 10.70102/afts.2025.1833.176
METAHEURISTIC-DRIVEN HYPERPARAMETER OPTIMIZATION FOR BERT IN SENTIMENT ANALYSIS
  • Oct 30, 2025
  • Archives for Technical Sciences
  • Alaa A El-Demerdash + 1 more

Sentiment analysis has come out as an important activity in natural language processing (NLP) applications whose data analysis is in high demand at present in the modern world. The BERT (Bidirectional Encoder Representations from Transformers) algorithm has proved to be extremely efficient when it comes to sentiment analysis tasks, and its potential is far exceeding that of conventional algorithms, unlocking their potential however would require fine tuning of their hyperparameters. It is quite a feat to optimise the BERT’s various hyperparameters due to the complicated interaction between them (e.g. the learning rate, batch size, dropout rate, attention heads). In this paper, the Salp Swarm Algorithm (SSA) is used as a bio-inspired metaheuristic optimization technique to optimize the fine-tuning process. Through SSA’s exceptionally efficient search capabilities in modelling multidimensional search space, BERT hyperparameters are optimized systematically to the sentiment classification tasks. A benchmark dataset for sentiment analysis (Sentiment140) is used to evaluate the proposed model. The novelty of the presented model is the fact that it dynamically adjusts its search behaviour in response to performance signals, thus it identifies better-performing parameter sets than conventional methods, leading to successful exploitation of the BERT algorithm that has produced high performing configurations. Extensive evaluations against 3 state-of-the-art search algorithms, namely manual tuning, grid search, and random search are conducted on the Sentiment140 benchmark dataset, demonstrating the superiority of the proposed SSA BERT optimization technique over state-of-the-art methods. The SSA-BERT model achieved a maximum accuracy of 96.4 percent, which is far better than manual tuning, grid search, and random search (65.0 percent, 69.5 percent and 72.0 percent respectively). It also performed better than other existing BERT models used in related literature, which showed accuracy levels between 46.4 and 75.7 percent in accordance with different benchmarks Sentiment analysis has come out as an important activity in natural language processing (NLP) applications whose data analysis is in high demand at present in the modern world. The BERT (Bidirectional Encoder Representations from Transformers) algorithm has proved to be extremely efficient when it comes to sentiment analysis tasks, and its potential is far exceeding that of conventional algorithms, unlocking their potential however would require fine tuning of their hyperparameters. It is quite a feat to optimise the BERT’s various hyperparameters due to the complicated interaction between them (e.g. the learning rate, batch size, dropout rate, attention heads). In this paper, the Salp Swarm Algorithm (SSA) is used as a bio-inspired metaheuristic optimization technique to optimize the fine-tuning process. Through SSA’s exceptionally efficient search capabilities in modelling multidimensional search space, BERT hyperparameters are optimized systematically to the sentiment classification tasks. A benchmark dataset for sentiment analysis (Sentiment140) is used to evaluate the proposed model. The novelty of the presented model is the fact that it dynamically adjusts its search behaviour in response to performance signals, thus it identifies better-performing parameter sets than conventional methods, leading to successful exploitation of the BERT algorithm that has produced high performing configurations. Extensive evaluations against 3 state-of-the-art search algorithms, namely manual tuning, grid search, and random search are conducted on the Sentiment140 benchmark dataset, demonstrating the superiority of the proposed SSA BERT optimization technique over state-of-the-art methods. The SSA-BERT model achieved a maximum accuracy of 96.4 percent, which is far better than manual tuning, grid search, and random search (65.0 percent, 69.5 percent and 72.0 percent respectively). It also performed better than other existing BERT models used in related literature, which showed accuracy levels between 46.4 and 75.7 percent in accordance with different benchmarks.

  • Research Article
  • 10.15294/7h63ma50
Sentiment Analysis on Twitter Social Media Regarding Covid-19 Vaccination with Naive Bayes Classifier (NBC) and Bidirectional Encoder Representations from Transformers (BERT)
  • Sep 30, 2024
  • Recursive Journal of Informatics
  • Angga Riski Dwi Saputra + 1 more

Abstract. The Covid-19 vaccine is an important tool to stop the Covid-19 pandemic, however, there are pros and cons from the public regarding this Covid-19 vaccine. Purpose: These responses were conveyed by the public in many ways, one of which is through social media such as Twitter. Responses given by the public regarding the Covid-19 vaccination can be analyzed and categorized into responses with positive, neutral or negative sentiments. Methods: In this study, sentiment analysis was carried out regarding Covid-19 vaccination originating from Twitter using the Naïve Bayes Classifier (NBC) and Bidirectional Encoder Representations from Transformers (BERT) algorithms. The data used in this study is public tweet data regarding the Covid-19 vaccination with a total of 29,447 tweet data in English. Result: Sentiment analysis begins with data preprocessing on the dataset used for data normalization and data cleaning before classification. Then word vectorization was performed with TF-IDF and data classification was performed using the Naïve Bayes Classifier (NBC) and Bidirectional Encoder Representations from Transformers (BERT) algorithms. From the classification results, an accuracy value of 73% was obtained for the Naïve Bayes Classifier (NBC) algorithm and 83% for the Bidirectional Encoder Representations from Transformers (BERT) algorithm. Novelty: A direct comparison between classical models such as NBC and modern deep learning models such as BERT offers new insights into the advantages and disadvantages of both approaches in processing Twitter data. Additionally, this study proposes temporal sentiment analysis, which allows evaluating changes in public sentiment regarding vaccination over time. Another innovation is the implementation of a hybrid approach to data cleansing that combines traditional methods with the natural language processing capabilities of BERT, which more effectively addresses typical Twitter data issues such as slang and spelling errors. Finally, this research also expands sentiment classification to be multi-label, identifying more specific sentiment categories such as trust, fear, or doubt, which provides a deeper understanding of public opinion.

  • Research Article
  • Cite Count Icon 1
  • 10.15294/rji.v2i2.67502
Sentiment Analysis on Twitter Social Media Regarding Covid-19 Vaccination with Naive Bayes Classifier (NBC) and Bidirectional Encoder Representations from Transformers (BERT)
  • Sep 30, 2024
  • Recursive Journal of Informatics
  • Angga Riski Dwi Saputra + 1 more

Abstract. The Covid-19 vaccine is an important tool to stop the Covid-19 pandemic, however, there are pros and cons from the public regarding this Covid-19 vaccine. Purpose: These responses were conveyed by the public in many ways, one of which is through social media such as Twitter. Responses given by the public regarding the Covid-19 vaccination can be analyzed and categorized into responses with positive, neutral or negative sentiments. Methods: In this study, sentiment analysis was carried out regarding Covid-19 vaccination originating from Twitter using the Naïve Bayes Classifier (NBC) and Bidirectional Encoder Representations from Transformers (BERT) algorithms. The data used in this study is public tweet data regarding the Covid-19 vaccination with a total of 29,447 tweet data in English. Result: Sentiment analysis begins with data preprocessing on the dataset used for data normalization and data cleaning before classification. Then word vectorization was performed with TF-IDF and data classification was performed using the Naïve Bayes Classifier (NBC) and Bidirectional Encoder Representations from Transformers (BERT) algorithms. From the classification results, an accuracy value of 73% was obtained for the Naïve Bayes Classifier (NBC) algorithm and 83% for the Bidirectional Encoder Representations from Transformers (BERT) algorithm. Novelty: A direct comparison between classical models such as NBC and modern deep learning models such as BERT offers new insights into the advantages and disadvantages of both approaches in processing Twitter data. Additionally, this study proposes temporal sentiment analysis, which allows evaluating changes in public sentiment regarding vaccination over time. Another innovation is the implementation of a hybrid approach to data cleansing that combines traditional methods with the natural language processing capabilities of BERT, which more effectively addresses typical Twitter data issues such as slang and spelling errors. Finally, this research also expands sentiment classification to be multi-label, identifying more specific sentiment categories such as trust, fear, or doubt, which provides a deeper understanding of public opinion.

  • PDF Download Icon
  • Research Article
  • Cite Count Icon 135
  • 10.3390/app13031445
Sentiment Analysis of Text Reviews Using Lexicon-Enhanced Bert Embedding (LeBERT) Model with Convolutional Neural Network
  • Jan 21, 2023
  • Applied Sciences
  • James Mutinda + 2 more

Sentiment analysis has become an important area of research in natural language processing. This technique has a wide range of applications, such as comprehending user preferences in ecommerce feedback portals, politics, and in governance. However, accurate sentiment analysis requires robust text representation techniques that can convert words into precise vectors that represent the input text. There are two categories of text representation techniques: lexicon-based techniques and machine learning-based techniques. From research, both techniques have limitations. For instance, pre-trained word embeddings, such as Word2Vec, Glove, and bidirectional encoder representations from transformers (BERT), generate vectors by considering word distances, similarities, and occurrences ignoring other aspects such as word sentiment orientation. Aiming at such limitations, this paper presents a sentiment classification model (named LeBERT) combining sentiment lexicon, N-grams, BERT, and CNN. In the model, sentiment lexicon, N-grams, and BERT are used to vectorize words selected from a section of the input text. CNN is used as the deep neural network classifier for feature mapping and giving the output sentiment class. The proposed model is evaluated on three public datasets, namely, Amazon products’ reviews, Imbd movies’ reviews, and Yelp restaurants’ reviews datasets. Accuracy, precision, and F-measure are used as the model performance metrics. The experimental results indicate that the proposed LeBERT model outperforms the existing state-of-the-art models, with a F-measure score of 88.73% in binary sentiment classification.

  • PDF Download Icon
  • Research Article
  • Cite Count Icon 16
  • 10.47813/2782-5280-2024-3-1-0311-0320
Bidirectional encoders to state-of-the-art: a review of BERT and its transformative impact on natural language processing
  • Mar 2, 2024
  • Информатика. Экономика. Управление - Informatics. Economics. Management
  • Rajesh Gupta

First developed in 2018 by Google researchers, Bidirectional Encoder Representations from Transformers (BERT) represents a breakthrough in natural language processing (NLP). BERT achieved state-of-the-art results across a range of NLP tasks while using a single transformer-based neural network architecture. This work reviews BERT's technical approach, performance when published, and significant research impact since release. We provide background on BERT's foundations like transformer encoders and transfer learning from universal language models. Core technical innovations include deeply bidirectional conditioning and a masked language modeling objective during BERT's unsupervised pretraining phase. For evaluation, BERT was fine-tuned and tested on eleven NLP tasks ranging from question answering to sentiment analysis via the GLUE benchmark, achieving new state-of-the-art results. Additionally, this work analyzes BERT's immense research influence as an accessible technique surpassing specialized models. BERT catalyzed adoption of pretraining and transfer learning for NLP. Quantitatively, over 10,000 papers have extended BERT and it is integrated widely across industry applications. Future directions based on BERT scale towards billions of parameters and multilingual representations. In summary, this work reviews the method, performance, impact and future outlook for BERT as a foundational NLP technique. We provide background on BERT's foundations like transformer encoders and transfer learning from universal language models. Core technical innovations include deeply bidirectional conditioning and a masked language modeling objective during BERT's unsupervised pretraining phase. For evaluation, BERT was fine-tuned and tested on eleven NLP tasks ranging from question answering to sentiment analysis via the GLUE benchmark, achieving new state-of-the-art results. Additionally, this work analyzes BERT's immense research influence as an accessible technique surpassing specialized models. BERT catalyzed adoption of pretraining and transfer learning for NLP. Quantitatively, over 10,000 papers have extended BERT and it is integrated widely across industry applications. Future directions based on BERT scale towards billions of parameters and multilingual representations. In summary, this work reviews the method, performance, impact and future outlook for BERT as a foundational NLP technique.

  • Research Article
  • 10.1177/20552076251393338
Ensemble learning for improved sentiment analysis in doctor–patient communication
  • May 1, 2025
  • Digital Health
  • Yufan Ge + 3 more

ObjectiveTo fill the benchmarking gap in clinician–patient sentiment analysis, we compare deep learning, transformer, and ensemble models for three-class (low/medium/high) sentiment classification in doctor–patient consultations.MethodsWe used a publicly available dataset of 3325 anonymized doctor–patient consultations from the Hugging Face repository (mahfoos/Patient-Doctor-Conversation) labeled as low, medium, or high severity. Preprocessing included text cleaning, tokenization, and padding; class balancing was applied only within the training split of each fold. Models evaluated were long short-term memory (LSTM), bidirectional LSTM (BiLSTM), convolutional neural networks (CNN), CNN–LSTM, and bidirectional encoder representations from transformers (BERT); an ensemble (hard voting over Logistic Regression, Random Forest, and Support Vector Classifier (SVC)) was also tested. Evaluation used stratified five-fold cross-validation, with metrics reported as mean ± SD across outer test folds (accuracy; macro-averaged precision/recall/F1). Interpretability was examined via BERT attention and feature attributions.ResultsThe ensemble achieved the highest accuracy (75.5 ± 0.5), outperforming BERT (66.98 ± 0.6), CNN–LSTM (65.68 ± 0.9), CNN (64.17 ± 0.8), BiLSTM (64.82 ± 0.7), and LSTM (58.66 ± 0.19). Class-wise analysis showed robust detection of high-severity interactions (e.g. ensemble F1 = 90.8 ± 1.3), while low-severity remained most challenging; the ensemble improved class 0 recall (58.7 ± 1.0), and BERT provided the highest class 0 precision (65.5 ± 1.0).ConclusionUnder stratified five-fold cross-validation, ensemble learning delivered the strongest and most balanced performance for three-class sentiment classification of clinician–patient dialogue, while transformers offered complementary precision on difficult cases. Attention- and feature-attribution analyses improved transparency, supporting clinical interpretability. Future work should scale to larger, multimodal (text/audio/vision) and multilingual datasets, and develop privacy-preserving, lightweight models for real-time deployment in clinical settings.

  • Conference Article
  • Cite Count Icon 4
  • 10.1109/icecaa55415.2022.9936448
Public Sentiment Assessment of Coronavirus-Specific Tweets using a Transformer-based BERT Classifier
  • Oct 13, 2022
  • Kanak Mahor + 1 more

Worldwide, the (COVID-19) pandemic had also affected people's daily routines. In general also during lockdown periods, people around the world use social media to express their thoughts and feelings about the epidemic which has interrupted their daily lives. There has been a huge spike in tweets about coronavirus on Twitter in a short period of time, including both positive and negative messages. As a result of the wide range of content in the tweets, the researchers have turned to sentiment analysis in order to gauge how the general public feels about COVID-19. According to the findings of this study, the best way to examine COVID-19 is to look at how people use Twitter to share their thoughts and opinions. Sentiment categorization can be accomplished by utilising a variety of feature sets as well as classifiers in combination with the suggested approach. Tweets collected from people with COVID-19 perceptions can be used to better understand and manage the epidemic. Positive, negative, as well as neutral emotion classifications are being used to classify tweets. In this study, Tweets containing specific information about the Coronavirus epidemic are used as sentiment analysis packages. Bidirectional Encoder Representations from Transformers (BERT) are used to identify sentiment categories, whereas the TF-IDF (term frequency-inverse document frequency) prototype is used to summarise the topics of postings. Trend analysis and qualitative methods are being used to identify negative sentiment traits. In general, when it comes to sentiment classification, the fine-tuned BERT is very accurate. In addition, the COVID-19-related post features of TF-IDF themes are accurately conveyed. Coronavirus tweet sentiments are analysed using a BERT and TF-IDF hybrid classifier. Single-sentence classification is transformed into pair-sentence classification, which solves BERT's performance issue in text classification problems. Our evaluation measures (accuracy= 0.70; precision= 0.67; recall= 0.64; and F1-score= 0.65) are used to evaluate the effectiveness of the classifier.

  • PDF Download Icon
  • Research Article
  • Cite Count Icon 12
  • 10.14569/ijacsa.2022.01312112
Aspect-based Sentiment Analysis for Bengali Text using Bidirectional Encoder Representations from Transformers (BERT)
  • Jan 1, 2022
  • International Journal of Advanced Computer Science and Applications
  • Moythry Manir Samia + 4 more

Public opinion is important for decision-making on numerous occasions for national growth in democratic countries like Bangladesh, the USA, and India. Sentiment analysis is a technique used to determine the polarity of opinions expressed in a text. The more complex stage of sentiment analysis is known as Aspect-Based Sentiment Analysis (ABSA), where it is possible to ascertain both the actual topics being discussed by the speakers as well as the polarity of each opinion. Nowadays, people leave comments on a variety of websites, including social networking sites, online news sources, and even YouTube video comment sections, on a wide range of topics. ABSA can play a significant role in utilizing these comments for a variety of objectives, including academic, commercial, and socioeconomic development. In English and many other popular European languages, there are many datasets for ABSA, but the Bengali language has very few of them. As a result, ABSA research on Bengali is relatively rare. In this paper, we present a Bengali dataset that has been manually annotated with five aspects and their corresponding sentiment. A baseline evaluation was also carried out using the Bidirectional Encoder Representations from Transformers (BERT) model, with 97% aspect detection accuracy and 77% sentiment classification accuracy. For aspect detection, the F1-score was 0.97 and for sentiment classification, it was 0.77.

  • Research Article
  • 10.21276/ierj24851626540690
TEXT BASED MASS OPINION MINING USING NEURAL NETWORK ALGORITHM WITH ABSTRACTIVE SUMMARIZATION
  • Apr 15, 2024
  • International Education and Research Journal
  • Chandhine V + 4 more

Our work, which focuses specifically on Twitter data, presents an advanced methodology for large-scale opinion mining. Sentiment analysis, or opinion mining, is the process of mechanically identifying and categorizing sentiments from textual data. We achieve accurate sentiment categorization and informative result summarizing by combining the latest neural network methods with abstractive summarization techniques. Tokenization, a crucial stage in natural language processing (NLP) activities, is where we start by using BERT (Bidirectional Encoder Representations from Transformers). By capturing the contextual subtleties of language used in Twitter tweets, BERT's contextual embeddings facilitate accurate tokenization. Our model successfully captures the textual data by utilizing BERT's capabilities, which paves the way for further sentiment analysis. We then apply Bidirectional Long Short-Term Memory (BiLSTM) to sentiment categorization. Recurrent neural networks (RNNs) such as BiLSTM are particularly good at recognizing sequential dependencies in data. Sentiment analysis relies heavily on word order since sentiments are frequently influenced by the context that words that come before and after give. Our model's incorporation of BiLSTM allows it to precisely classify tweets as either positive or negative by capturing these sequential dependencies. We also present an abstractive summary element to produce brief summaries that capture dominant attitudes in the dataset. Abstractive summarization generates logical summaries that encapsulate the main ideas of numerous tweets, going beyond simple sentence selection and concatenation. We evaluate the quality of the generated summaries and the accuracy of sentiment categorization achieved by our model through extensive experimentation on a representative and diversified Twitter dataset. Our findings show how reliable and effective our approach is at identifying significant sentiment trends and offering insightful information about the general thoughts shared on Twitter. In the end, our work advances the field of opinion mining and provides useful instruments for deciphering and evaluating massive amounts of social media data.

  • Research Article
  • Cite Count Icon 37
  • 10.1108/jhtt-09-2021-0259
Discovering a tourism destination with social media data: BERT-based sentiment analysis
  • Aug 19, 2022
  • Journal of Hospitality and Tourism Technology
  • Marlon Santiago Viñán-Ludeña + 1 more

使用社交媒体数据发现旅游目的地:基于 bert 的情感分析。研究目的这项工作的主要目的是使用情感分析技术和来自 Twitter 和 Instagram 的数据来分析旅游目的地, 以便找到最具代表性的实体(或地点)和用户的感知(或方面)。研究设计/方法/途径我们使用 90,725 个 Instagram 帖子和 235,755 个 Twitter 推文来分析格拉纳达(西班牙)的旅游业, 以确定旅行者在两个社交媒体网站上提到的重要地点和看法。我们使用了几种方法对英语和西班牙语文本进行情感分类, 包括深度学习模型。研究发现测试集中的最佳结果是使用来自Transformers (BERT) 模型的双向编码器表示 (BERT) 用于西班牙语文本和Tweeteval 用于英语文本, 这些结果随后用于分析我们的数据集。然后可以确定最重要的实体和方面, 这反过来又为研究人员、从业人员、旅行者和旅游管理者提供了有趣的见解, 从而可以改进服务并制定更好的营销策略。研究局限性我们提出了一个用于执行情感分类的西班牙旅游 BERT 模型, 以及通过主题标签找到地点并揭示每个地点的重要负面方面的过程。实践意义该研究使管理人员和从业人员能够使用我们发布的西班牙旅游数据集实施西班牙-BERT 模型, 以便在应用程序中采用该数据集, 以找到正面和负面的看法。研究原创性本研究提出了一种如何在旅游领域应用情感分析的新方法。首先, 介绍了评估不同现有模型和工具的方法; 其次, 使用 BERT(深度学习模型)训练模型; 第三, 提出了如何通过标签识别目的地地点的接受度的方法, 最后通过实体和方面的识别来评估用户表达积极性(消极性)的原因。

  • Research Article
  • Cite Count Icon 44
  • 10.11591/ijeecs.v25.i2.pp1131-1139
Crime prediction using a hybrid sentiment analysis approach based on the bidirectional encoder representations from transformers
  • Feb 1, 2022
  • Indonesian Journal of Electrical Engineering and Computer Science
  • Mohammed Boukabous + 1 more

Sentiment analysis (SA) is widely used today in many areas such as crime detection (security intelligence) to detect potential security threats in realtime using social media platforms such as Twitter. The most promising techniques in sentiment analysis are those of deep learning (DL), particularly bidirectional encoder representations from transformers (BERT) in the field of natural language processing (NLP). However, employing the BERT algorithm to detect crimes requires a crime dataset labeled by the lexiconbased approach. In this paper, we used a hybrid approach that combines both lexicon-based and deep learning, with BERT as the DL model. We employed the lexicon-based approach to label our Twitter dataset with a set of normal and crime-related lexicons; then, we used the obtained labeled dataset to train our BERT model. The experimental results show that our hybrid technique outperforms existing approaches in several metrics, with 94.91% and 94.92% in accuracy and F1-score respectively.

  • Research Article
  • Cite Count Icon 12
  • 10.14569/ijacsa.2024.0151069
Combining BERT and CNN for Sentiment Analysis A Case Study on COVID-19
  • Jan 1, 2024
  • International Journal of Advanced Computer Science and Applications
  • Gunjan Kumar + 7 more

This research focuses on sentiment analysis to understand public opinion on various topics, with an emphasis on COVID-19 discussions on Twitter. By utilizing state-of-the-art Machine Learning (ML) and Natural Language Processing (NLP) techniques, the study analyzes sentiment data to provide valuable insights. The process begins with data preparation, involving text cleaning and length filtering to optimize the dataset for analysis. Two models are employed: a Bidirectional Encoder Representations from Transformers (BERT)-based Deep Learning (DL) model and a Convolutional Neural Network (CNN). The BERT model leverages transfer learning, demonstrating strong performance in sentiment classification, while the CNN model excels at extracting contextual features from the input text. To further enhance accuracy, an ensemble model integrates predictions from both approaches. The study emphasizes the ensemble technique’s value for more precise sentiment analysis. Evaluation metrics, including accuracy, classification reports, and confusion matrices, validate the effectiveness of the proposed models and the ensemble approach. This research contributes to the growing field of social media sentiment analysis, particularly during global health crises like COVID-19, and underscores its potential to aid informed decision-making based on public sentiment.

  • Research Article
  • 10.37394/232032.2024.2.15
Financial Report Sentiment Analysis Using Loughran-mcdonald Dictionary and BERT
  • Jun 24, 2024
  • Financial Engineering
  • Sheetal R + 1 more

In the ever-changing world of financial markets, understanding investor behavior and making informed decisions relies heavily on sentiment analysis. This study delves into the integration of traditional techniques, such as the Loughran- McDonald dictionary, with advanced natural language processing (NLP) methods utilizing BERT (Bidirectional Encoder Representations from Transformers). The goal is to enhance the accuracy and depth of sentiment analysis in financial reports.To begin, we employ the specialized Loughran-McDonald dictionary designed for financial sentiment analysis. This lexicon includes domainspecific word lists for positive and negative sentiments, forming a solid foundation for sentiment scoring. Expanding on this foundation, we incorporate BERT, an advanced transformerbased NLP model. BERT’s contextual understanding of language and ability to capture intricate semantic relationships within financial texts aim to overcome the limitations of rule-based sentiment analysis. The methodology involves preprocessing financial reports, integrating Loughran-McDonald sentiment scores, and fine-tuning BERT for financial sentiment classification. This hybrid approach leverages both the domain expertise encoded in the dictionary and BERT’s contextual comprehension of financial jargon and nuances. We validate and evaluate our implementation using a diverse dataset comprising quarterly earnings releases, annual reports, and other relevant disclosures. Performance metrics such as precision, recall, and F1 score are analyzed to assess the effectiveness of our hybrid approach compared to individual methods. The findings have significant implications for financial analysts, investors, and policymakers by providing a more nuanced understanding of sentiment in financial reports. Our hybrid approach aims to offer improved accuracy in capturing sentiment polarity while facilitating more informed decision-making in today’s complex and dynamic realm of financial markets.

  • PDF Download Icon
  • Research Article
  • Cite Count Icon 17
  • 10.1007/s11227-023-05319-8
Vaccine sentiment analysis using BERT + NBSVM and geo-spatial approaches
  • May 7, 2023
  • The Journal of Supercomputing
  • Areeba Umair + 2 more

Since the spread of the coronavirus flu in 2019 (hereafter referred to as COVID-19), millions of people worldwide have been affected by the pandemic, which has significantly impacted our habits in various ways. In order to eradicate the disease, a great help came from unprecedentedly fast vaccines development along with strict preventive measures adoption like lockdown. Thus, world wide provisioning of vaccines was crucial in order to achieve the maximum immunization of population. However, the fast development of vaccines, driven by the urge of limiting the pandemic caused skeptical reactions by a vast amount of population. More specifically, the people’s hesitancy in getting vaccinated was an additional obstacle in fighting COVID-19. To ameliorate this scenario, it is important to understand people’s sentiments about vaccines in order to take proper actions to better inform the population. As a matter of fact, people continuously update their feelings and sentiments on social media, thus a proper analysis of those opinions is an important challenge for providing proper information to avoid misinformation. More in detail, sentiment analysis (Wankhade et al. in Artif Intell Rev 55(7):5731–5780, 2022. https://doi.org/10.1007/s10462-022-10144-1) is a powerful technique in natural language processing that enables the identification and classification of people feelings (mainly) in text data. It involves the use of machine learning algorithms and other computational techniques to analyze large volumes of text and determine whether they express positive, negative or neutral sentiment. Sentiment analysis is widely used in industries such as marketing, customer service, and healthcare, among others, to gain actionable insights from customer feedback, social media posts, and other forms of unstructured textual data. In this paper, Sentiment Analysis will be used to elaborate on people reaction to COVID-19 vaccines in order to provide useful insights to improve the correct understanding of their correct usage and possible advantages. In this paper, a framework that leverages artificial intelligence (AI) methods is proposed for classifying tweets based on their polarity values. We analyzed Twitter data related to COVID-19 vaccines after the most appropriate pre-processing on them. More specifically, we identified the word-cloud of negative, positive, and neutral words using an artificial intelligence tool to determine the sentiment of tweets. After this pre-processing step, we performed classification using the BERT + NBSVM model to classify people’s sentiments about vaccines. The reason for choosing to combine bidirectional encoder representations from transformers (BERT) and Naive Bayes and support vector machine (NBSVM ) can be understood by considering the limitation of BERT-based approaches, which only leverage encoder layers, resulting in lower performance on short texts like the ones used in our analysis. Such a limitation can be ameliorated by using Naive Bayes and Support Vector Machine approaches that are able to achieve higher performance in short text sentiment analysis. Thus, we took advantage of both BERT features and NBSVM features to define a flexible framework for our sentiment analysis goal related to vaccine sentiment identification. Moreover, we enrich our results with spatial analysis of the data by using geo-coding, visualization, and spatial correlation analysis to suggest the most suitable vaccination centers to users based on the sentiment analysis outcomes. In principle, we do not need to implement a distributed architecture to run our experiments as the available public data are not massive. However, we discuss a high-performance architecture that will be used if the collected data scales up dramatically. We compared our approach with the state-of-art methods by comparing most widely used metrics like Accuracy, Precision, Recall and F-measure. The proposed BERT + NBSVM outperformed alternative models by achieving 73% accuracy, 71% precision, 88% recall and 73% F-measure for classification of positive sentiments while 73% accuracy, 71% precision, 74% recall and 73% F-measure for classification of negative sentiments respectively. These promising results will be properly discussed in next sections. The use of artificial intelligence methods and social media analysis can lead to a better understanding of people’s reactions and opinions about any trending topic. However, in the case of health-related topics like COVID-19 vaccines, proper sentiment identification could be crucial for implementing public health policies. More in detail, the availability of useful findings on user opinions about vaccines can help policymakers design proper strategies and implement ad-hoc vaccination protocols according to people’s feelings, in order to provide better public service. To this end, we leveraged geospatial information to support effective recommendations for vaccination centers.

  • Research Article
  • Cite Count Icon 1
  • 10.11591/ijeecs.v34.i1.pp497-507
Sentiment analysis and classification of Ghanaian football tweets from the 2022 FIFA World Cup
  • Apr 1, 2024
  • Indonesian Journal of Electrical Engineering and Computer Science
  • Eshun Michael + 5 more

Football as an attractive sport generates huge volumes of tweets concerning fans' opinions, feelings, and judgments during prime events. Such data can be leveraged in sentiment analysis, an algorithmic approach to analyzing text in tweets by extracting emotional tones. This paper presents a novel benchmark dataset of 132,115 tweets collected during the 2022 world cup on 𝕏 (formerly Twitter) for football-related sentiment classification. We also performed sentiment analysis on the dataset using lexicon-based tools, traditional machine learning algorithms, and pre-trained models, robustly optimized bidirectional encoder representations from transformers (BERT)- pretraining approach RoBERTa and distilled version of BERT (DistilBERT) to understand the emotions and reactions of football fans during different phases of the football matches. Results from the study indicate that most tweets had neutral sentiments in both context-aware and context-free analysis. We also describe our novel GhaFootBERT, a sentiment classification model based on transfer learning on BERT, which provides an effective approach to sentiment classification of football-related tweets. Our model performs robustly, outperforming the traditional models with 92% accuracy.

Save Icon
Up Arrow
Open/Close
Notes

Save Important notes in documents

Highlight text to save as a note, or write notes directly

You can also access these Documents in Paperpal, our AI writing tool

Powered by our AI Writing Assistant