B-TTDb: A Database of Turkish Tweets for Predicting the Top One Hundred Emojis

Yiltan Bitirim

doi:10.1145/3681783

Abstract

Emoji prediction is an important research task that focuses on finding the most appropriate emoji(s) quickly and effortlessly for a specific text. Now that Turkish is on the list of the top 20 most spoken languages in the world and there are a considerable number of Turkish-speaking social media users, studying emoji prediction in Turkish holds significant value. In this study, a Turkish tweets database, named Bitirim's Turkish Tweets Database (B-TTDb), was constructed for academic and industrial studies based on the prediction of the top 100 emojis. B-TTDb consists of four datasets. The first dataset includes raw tweets, the second dataset is the organized version of the first dataset, the third dataset is the pre-processed version of the second dataset, and the last one is the organized version of the third dataset. The last one is the final version and it is named Bitirim's Dataset (B-D). It includes a total of 158,201 unique tweets belonging to the top 100 emoji classes. For database validation, experiments were conducted on B-D with popular machine learning algorithms for the top 10, 20, 50, and 100 emojis. This study could be considered as the first study that contributes to the literature by the first validated large database of Turkish tweets that includes such a large number of emojis. In addition, B-TTDb could be a basis as well as motivation for various further studies.

Talk to us

Join us for a 30 min session where you can share your feedback and ask us any queries you have

Schedule a call

R Discovery Prime

R Discovery Prime

B-TTDb: A Database of Turkish Tweets for Predicting the Top One Hundred Emojis

Abstract

Talk to us

Similar Papers

More From: ACM Transactions on the Web

Lead the way for us

Journal: ACM Transactions on the Web	Publication Date: Jul 24, 2024
License type: cc-by

Similar Papers

Emoji Identification and Prediction in Hebrew Political Corpus
-
-
--
01 Jan 2019
01 Jan 2019

Emoji Identification and Prediction in Hebrew Political Corpus
Chaya Liebeskind
Issues in Informing Science and Information Technology | VOL. 16
Chaya LiebeskindChaya Liebeskind
01 Jan 2019
Issues in Informing Science and Information Technology | VOL. 16

International Spinal Cord Injury Core Data Set (version 2.0)-including standardization of reporting.
F Biering-Sørensen ... S Charlifue
Spinal Cord | VOL. 55
F Biering-Sørensen, et. al.F Biering-Sørensen ... S Charlifue
30 May 2017
Spinal Cord | VOL. 55

Sentiment analysis on Twitter: A text mining approach to the Syrian refugee crisis
Nazan Öztürk ... Serkan Ayvaz
Telematics and Informatics | VOL. 35
Nazan Öztürk, et. al.Nazan Öztürk ... Serkan Ayvaz
23 Oct 2017
Telematics and Informatics | VOL. 35

Editage

Paperpal

R Discovery

Mind the Graph

R Discovery Prime

R Discovery Prime

B-TTDb: A Database of Turkish Tweets for Predicting the Top One Hundred Emojis

Abstract

Talk to us

Similar Papers

More From: ACM Transactions on the Web