An empiric validation of linguistic features in machine learning models for fake news detection

Eduardo Puraivan,René Venegas,Fabián Riquelme

doi:10.1016/j.datak.2023.102207

Abstract

The diffusion of fake news is a growing problem with a high and negative social impact. There are several approaches to address the detection of fake news. This work focuses on a hybrid approach based on functional linguistic features and machine learning. There are several recent works with this approach. However, there are no clear guidelines on which linguistic features are most appropriate nor how to justify their use. Furthermore, many classification results are modest compared to recent advances in natural language processing. Our proposal considers 88 features organized in surface information, part of speech, discursive characteristics, and readability indices. On a 42 677 news database, we show that the classification results outperform previous work, even outperforming state-of-the-art techniques such as BERT, reaching 99.99% accuracy. A proper selection of linguistic features is crucial for interpretability as well as the performance of the models. In this sense, our proposal contributes to the intentional selection of linguistic features, overcoming current technical issues. We identified 32 features that show differences between the type of news. The results are highly competitive in the classification and simple to implement and interpret.

Talk to us

Join us for a 30 min session where you can share your feedback and ask us any queries you have

Schedule a call

R Discovery Prime

R Discovery Prime

An empiric validation of linguistic features in machine learning models for fake news detection

Abstract

Talk to us

Similar Papers

More From: Data & Knowledge Engineering

Lead the way for us

Journal: Data & Knowledge Engineering	Publication Date: Aug 2, 2023
Citations: 1

Similar Papers

Linguistic analysis of datasets for semantic textual similarity
Chunlin Wang ... Elisabet Comelles
Digital Scholarship in the Humanities | VOL. 35
Chunlin Wang, et. al.Chunlin Wang ... Elisabet Comelles
27 Apr 2019
Digital Scholarship in the Humanities | VOL. 35

Fake News Detection Using A Deep Neural Network
Rohit Kumar Kaliyar
-
Rohit Kumar KaliyarRohit Kumar Kaliyar
01 Dec 2018
01 Dec 2018

Generating Abstract Test Cases from User Requirements using MDSE and NLP
Sai Chaithra Allala ... Tariq M King
-
Sai Chaithra Allala, et. al.Sai Chaithra Allala ... Tariq M King
01 Dec 2022
01 Dec 2022

Natural Language Processing with Optimal Deep Learning Based Fake News Classification
Sara A Althubiti ... Fayadh Alenezi
Computers, Materials & Continua | VOL. 73
Sara A Althubiti, et. al.Sara A Althubiti ... Fayadh Alenezi
01 Jan 2021
Computers, Materials & Continua | VOL. 73

Editage

Paperpal

R Discovery

Mind the Graph

R Discovery Prime

R Discovery Prime

An empiric validation of linguistic features in machine learning models for fake news detection

Abstract

Talk to us

Similar Papers

More From: Data &amp; Knowledge Engineering

More From: Data & Knowledge Engineering