The Application of Random Forest to the Classification of Fake News

Najwan Thair Ali,Karrar Falih Hassan,Muataz Najim Abdullah,Zainab Salam Al-Hchimy,N Aldahan,A.J Ramadhan

doi:10.1051/bioconf/20249700049

Abstract

Fake News is one of the most widespread phenomenon with significant consequences on our daily life, particularly in the political realm. Due to the increasing use of the internet and social media, it is now much simpler to propagate false information. Therefore, the identification of elusive news is a significant issue that must be addressed, mostly due to obstacles such as the limited number of benchmark datasets and the volume of news produced per second. This study suggested using comparative data analysis based on random forest machine learning algorithm to identify bogus 4news. In this study the size of the whole dataset is 20.761 fake news record, whereas the size of it is 4.345 records. The first step in the data preparation process is to remove any unnecessary special characters, numbers, English letters, and whitespace. Before implementing the proposed classification algorithms, the most prevalent feature extraction approach (TF-IDF) is used. The data indicate that the highest level of accuracy attained was 88.24%.

Full Text