Unstructured Text Data Research Articles

The metaverse has become one of the most popular concepts of recent times. Companies and entrepreneurs are fiercely competing to invest and take part in this virtual world. Millions of people globally are anticipated to spend much of their time in the metaverse, regardless of their age, gender, ethnicity, or culture. There are few comprehensive studies on the positive/negative sentiment and effect of the newly identified, but not well defined, metaverse concept that is already fast evolving the digital landscape. Thereby, this study aimed to better understand the metaverse concept, by, firstly, identifying the positive and negative sentiment characteristics and, secondly, by revealing the associations between the metaverse concept and other related concepts. To do so, this study used Natural Language Processing (NLP) methods, specifically Artificial Intelligence (AI) with computational qualitative analysis. The data comprised metaverse articles from 2021 to 2022 published on The Guardian website, a key global mainstream media outlet. To perform thematic content analysis of the qualitative data, this research used the Leximancer software, and the The Natural Language Toolkit (NLTK) from NLP libraries were used to identify sentiment. Further, an AI-based Monkeylearn API was used to make sectoral classifications of the main topics that emerged in the Leximancer analysis. The key themes which emerged in the Leximancer analysis, included "metaverse", "Facebook", "games" and "platforms". The sentiment analysis revealed that of all articles published in the period of 2021–2022 about the metaverse, 61% (n = 622) were positive, 30% (n = 311) were negative, and 9% (n = 90) were neutral. Positive discourses about the metaverse were found to concern key innovations that the virtual experiences brought to users and companies with the support of the technological infrastructure of blockchain, algorithms, NFTs, led by the gaming world. Negative discourse was found to evidence various problems (misinformation, harmful content, algorithms, data, and equipment) that occur during the use of Facebook and other social media platforms, and that individuals encountered harm in the metaverse or that the metaverse produces new problems. Monkeylearn findings revealed “marketing/advertising/PR” role, “Recreational” business, “Science & Technology” events as the key content topics. This study’s contribution is twofold: first, it showcases a novel way to triangulate qualitative data analysis of large unstructured textual data as a method in exploring the metaverse concept; and second, the study reveals the characteristics of the metaverse as a concept, as well as its association with other related concepts. Given that the topic of the metaverse is new, this is the first study, to our knowledge, to do both.

Read full abstract

Since the spread of the coronavirus flu in 2019 (hereafter referred to as COVID-19), millions of people worldwide have been affected by the pandemic, which has significantly impacted our habits in various ways. In order to eradicate the disease, a great help came from unprecedentedly fast vaccines development along with strict preventive measures adoption like lockdown. Thus, world wide provisioning of vaccines was crucial in order to achieve the maximum immunization of population. However, the fast development of vaccines, driven by the urge of limiting the pandemic caused skeptical reactions by a vast amount of population. More specifically, the people’s hesitancy in getting vaccinated was an additional obstacle in fighting COVID-19. To ameliorate this scenario, it is important to understand people’s sentiments about vaccines in order to take proper actions to better inform the population. As a matter of fact, people continuously update their feelings and sentiments on social media, thus a proper analysis of those opinions is an important challenge for providing proper information to avoid misinformation. More in detail, sentiment analysis (Wankhade et al. in Artif Intell Rev 55(7):5731–5780, 2022. https://doi.org/10.1007/s10462-022-10144-1) is a powerful technique in natural language processing that enables the identification and classification of people feelings (mainly) in text data. It involves the use of machine learning algorithms and other computational techniques to analyze large volumes of text and determine whether they express positive, negative or neutral sentiment. Sentiment analysis is widely used in industries such as marketing, customer service, and healthcare, among others, to gain actionable insights from customer feedback, social media posts, and other forms of unstructured textual data. In this paper, Sentiment Analysis will be used to elaborate on people reaction to COVID-19 vaccines in order to provide useful insights to improve the correct understanding of their correct usage and possible advantages. In this paper, a framework that leverages artificial intelligence (AI) methods is proposed for classifying tweets based on their polarity values. We analyzed Twitter data related to COVID-19 vaccines after the most appropriate pre-processing on them. More specifically, we identified the word-cloud of negative, positive, and neutral words using an artificial intelligence tool to determine the sentiment of tweets. After this pre-processing step, we performed classification using the BERT + NBSVM model to classify people’s sentiments about vaccines. The reason for choosing to combine bidirectional encoder representations from transformers (BERT) and Naive Bayes and support vector machine (NBSVM ) can be understood by considering the limitation of BERT-based approaches, which only leverage encoder layers, resulting in lower performance on short texts like the ones used in our analysis. Such a limitation can be ameliorated by using Naive Bayes and Support Vector Machine approaches that are able to achieve higher performance in short text sentiment analysis. Thus, we took advantage of both BERT features and NBSVM features to define a flexible framework for our sentiment analysis goal related to vaccine sentiment identification. Moreover, we enrich our results with spatial analysis of the data by using geo-coding, visualization, and spatial correlation analysis to suggest the most suitable vaccination centers to users based on the sentiment analysis outcomes. In principle, we do not need to implement a distributed architecture to run our experiments as the available public data are not massive. However, we discuss a high-performance architecture that will be used if the collected data scales up dramatically. We compared our approach with the state-of-art methods by comparing most widely used metrics like Accuracy, Precision, Recall and F-measure. The proposed BERT + NBSVM outperformed alternative models by achieving 73% accuracy, 71% precision, 88% recall and 73% F-measure for classification of positive sentiments while 73% accuracy, 71% precision, 74% recall and 73% F-measure for classification of negative sentiments respectively. These promising results will be properly discussed in next sections. The use of artificial intelligence methods and social media analysis can lead to a better understanding of people’s reactions and opinions about any trending topic. However, in the case of health-related topics like COVID-19 vaccines, proper sentiment identification could be crucial for implementing public health policies. More in detail, the availability of useful findings on user opinions about vaccines can help policymakers design proper strategies and implement ad-hoc vaccination protocols according to people’s feelings, in order to provide better public service. To this end, we leveraged geospatial information to support effective recommendations for vaccination centers.

Read full abstract

Unstructured Text Data Research Articles

Related Topics

Articles published on Unstructured Text Data

Text Classification System Using Text Mining with XGBoost Method

Role of Text Mining in Extracting Valuable Information from Text Data

Development of an artificial intelligence bacteremia prediction model and evaluation of its impact on physician predictions focusing on uncertainty

Business Analytics Using Predictive Algorithms

A survey of text detection and recognition algorithms based on deep learning technology

Sentiment based emotion classification in unstructured textual data using dual stage deep model

Topological properties and organizing principles of semantic networks

Text Document Clustering Approach by Improved Sine Cosine Algorithm

Online maintenance of evolving knowledge graphs with RDFS-based saturation and why-provenance support

Behavior-Based Anomaly Detection in Log Data of Physical Access Control Systems

Classification of Unstructured Customer Complaint Text Data for Potential Vehicle Defect Detection

Exploring Artificial Intelligence in English Language Arts with StoryQ

Exploring Latent Themes-Analysis of various Topic Modelling Algorithms

Mobilizing Text As Data

Applications of Natural Language Processing to Geoscience Text Data and Prospectivity Modeling

Factors Considered Important by Healthcare Professionals for the Management of Using Complementary Therapy in Diabetes: A Text-Mining Analysis.

An exploratory content and sentiment analysis of the guardian metaverse articles using leximancer and natural language processing

A case study of using natural language processing to extract consumer insights from tweets in American cities for public health crises

Large-Scale Knowledge Synthesis and Complex Information Retrieval from Biomedical Documents

Vaccine sentiment analysis using BERT + NBSVM and geo-spatial approaches

Lead the way for us

Editage

Paperpal

R Discovery

Mind the Graph

Unstructured Text Data Research Articles

Related Topics

Articles published on Unstructured Text Data

Text Classification System Using Text Mining with XGBoost Method

Role of Text Mining in Extracting Valuable Information from Text Data

Development of an artificial intelligence bacteremia prediction model and evaluation of its impact on physician predictions focusing on uncertainty

Business Analytics Using Predictive Algorithms

A survey of text detection and recognition algorithms based on deep learning technology

Sentiment based emotion classification in unstructured textual data using dual stage deep model

Topological properties and organizing principles of semantic networks

Text Document Clustering Approach by Improved Sine Cosine Algorithm

Online maintenance of evolving knowledge graphs with RDFS-based saturation and why-provenance support

Behavior-Based Anomaly Detection in Log Data of Physical Access Control Systems

Classification of Unstructured Customer Complaint Text Data for Potential Vehicle Defect Detection

Exploring Artificial Intelligence in English Language Arts with StoryQ

Exploring Latent Themes-Analysis of various Topic Modelling Algorithms

Mobilizing Text As Data

Applications of Natural Language Processing to Geoscience Text Data and Prospectivity Modeling

Factors Considered Important by Healthcare Professionals for the Management of Using Complementary Therapy in Diabetes: A Text-Mining Analysis.

An exploratory content and sentiment analysis of the guardian metaverse articles using leximancer and natural language processing

A case study of using natural language processing to extract consumer insights from tweets in American cities for public health crises

Large-Scale Knowledge Synthesis and Complex Information Retrieval from Biomedical Documents

Vaccine sentiment analysis using BERT + NBSVM and geo-spatial approaches