Social Media Mining Toolkit (SMMT).

Ramya Tekumalla,Juan M Banda

doi:10.5808/gi.2020.18.2.e16

Abstract

There has been a dramatic increase in the popularity of utilizing social media data for research purposes within the biomedical community. In PubMed alone, there have been nearly 2,500 publication entries since 2014 that deal with analyzing social media data from Twitter and Reddit. However, the vast majority of those works do not share their code or data for replicating their studies. With minimal exceptions, the few that do, place the burden on the researcher to figure out how to fetch the data, how to best format their data, and how to create automatic and manual annotations on the acquired data. In order to address this pressing issue, we introduce the Social Media Mining Toolkit (SMMT), a suite of tools aimed to encapsulate the cumbersome details of acquiring, preprocessing, annotating and standardizing social media data. The purpose of our toolkit is for researchers to focus on answering research questions, and not the technical aspects of using social media data. By using a standard toolkit, researchers will be able to acquire, use, and release data in a consistent way that is transparent for everybody using the toolkit, hence, simplifying research reproducibility and accessibility in the social media domain.

Highlights

In the last six years, there has been a great influx of research works that describe different types of research works using Twitter and Reddit data, nearly 2,500 papers are found in PubMed [1]
In an attempt to shift the biomedical community into better practices for research transparency and reproducibility, we introduce the Social Media Mining Toolkit (SMMT), a suite of tools aimed to encapsulate the cumbersome details of acquiring, preprocessing, annotating, and standardizing social media data
Parallel to tools like SMMT, there are other research groups that are outlining frameworks to streamline the mining of social media like Sarker et al [12], which are complementary to the use and need of this tool

Summary

Introduction

In the last six years, there has been a great influx of research works that describe different types of research works using Twitter and Reddit data, nearly 2,500 papers are found in PubMed [1]. The data acquisition methodology is different on each study and seldomly reported, a crucial step towards reproducibility of any of their analyses. When it comes to using Twitter data for drug identification and pharmacovigilance tasks, authors of works like [7,8,9] have been consistently releasing publicly available datasets, software tools, and complete Natural Language Processing (NLP) systems with their works. Parallel to tools like SMMT, there are other research groups that are outlining frameworks to streamline the mining of social media like Sarker et al [12], which are complementary to the use and need of this tool

Methods

Findings

Discussion

Full Text

Paper version not known

Open DOI Link

Talk to us

Join us for a 30 min session where you can share your feedback and ask us any queries you have

Schedule a call

Journal: Genomics & Informatics	Publication Date: Jun 15, 2020
Citations: 43	License type: CC BY 4.0

R Discovery Prime

R Discovery Prime

Social Media Mining Toolkit (SMMT).

Abstract

Highlights

Summary

Talk to us

Similar Papers

More From: Genomics & Informatics

Lead the way for us

Similar Papers

An ensemble heterogeneous classification methodology for discovering health-related knowledge in social media messages
Suppawong Tuarob ... Nilam Ram
Journal of Biomedical Informatics | VOL. 49
Suppawong Tuarob, et. al.Suppawong Tuarob ... Nilam Ram
16 Mar 2014
Journal of Biomedical Informatics | VOL. 49

A Pilot Study of the Application of Natural Language Interpretation to Identify Insights on Acute Lymphocytic Leukemia and Its Treatment from Social Media Data
Rebecca Crawford ... Cheryl Truchan
Blood | VOL. 138
Rebecca Crawford, et. al.Rebecca Crawford ... Cheryl Truchan
05 Nov 2021
Blood | VOL. 138

FluMapper: A cyberGIS application for interactive analysis of massive location‐based social media
Anand Padmanabhan ... Kiumars Soltani
Concurrency and Computation: Practice and Experience | VOL. 26
Anand Padmanabhan, et. al.Anand Padmanabhan ... Kiumars Soltani
16 May 2014
Concurrency and Computation: Practice and Experience | VOL. 26

Breast Cancer Symptom Clusters Derived from Social Media and Research Study Data Using Improved K-Medoid Clustering.
Qing Ping ... Sarah A Marshall
IEEE Transactions on Computational Social Systems | VOL. 3
Qing Ping, et. al.Qing Ping ... Sarah A Marshall
01 Jun 2016
IEEE Transactions on Computational Social Systems | VOL. 3

Editage

Paperpal

R Discovery

Mind the Graph

R Discovery Prime

R Discovery Prime

Social Media Mining Toolkit (SMMT).

Abstract

Highlights

Summary

Talk to us

Similar Papers

More From: Genomics &amp; Informatics

More From: Genomics & Informatics