Multilingual Hate Speech Detection: A Semi-Supervised Generative Adversarial Approach.

Khouloud Mnassri,Reza Farahbakhsh,Noel Crespi

doi:10.3390/e26040344

Abstract

Social media platforms have surpassed cultural and linguistic boundaries, thus enabling online communication worldwide. However, the expanded use of various languages has intensified the challenge of online detection of hate speech content. Despite the release of multiple Natural Language Processing (NLP) solutions implementing cutting-edge machine learning techniques, the scarcity of data, especially labeled data, remains a considerable obstacle, which further requires the use of semisupervised approaches along with Generative Artificial Intelligence (Generative AI) techniques. This paper introduces an innovative approach, a multilingual semisupervised model combining Generative Adversarial Networks (GANs) and Pretrained Language Models (PLMs), more precisely mBERT and XLM-RoBERTa. Our approach proves its effectiveness in the detection of hate speech and offensive language in Indo-European languages (in English, German, and Hindi) when employing only 20% annotated data from the HASOC2019 dataset, thereby presenting significantly high performances in each of multilingual, zero-shot crosslingual, and monolingual training scenarios. Our study provides a robust mBERT-based semisupervised GAN model (SS-GAN-mBERT) that outperformed the XLM-RoBERTa-based model (SS-GAN-XLM) and reached an average F1 score boost of 9.23% and an accuracy increase of 5.75% over the baseline semisupervised mBERT model.

Talk to us

Join us for a 30 min session where you can share your feedback and ask us any queries you have

Schedule a call

Journal: Entropy	Publication Date: Apr 18, 2024
Citations: 1	License type: CC BY 4.0

R Discovery Prime

R Discovery Prime

Multilingual Hate Speech Detection: A Semi-Supervised Generative Adversarial Approach.

Abstract

Talk to us

Similar Papers

More From: Entropy

Lead the way for us

Similar Papers

Hate speech detection: A comprehensive review of recent works
Ankita Gandhi ... Amir Hussain
Expert Systems | VOL. 41
Ankita Gandhi, et. al.Ankita Gandhi ... Amir Hussain
25 Feb 2024
Expert Systems | VOL. 41

Hate speech detection in low-resourced Indian languages: An analysis of transformer-based monolingual and multilingual models with cross-lingual experiments
Koyel Ghosh ... Apurbalal Senapati
Natural Language Processing | VOL. -
Koyel Ghosh, et. al.Koyel Ghosh ... Apurbalal Senapati
27 Aug 2024
Natural Language Processing | VOL. -

Towards Automatic Detection and Explanation of Hate Speech and Offensive Language
Wyatt Dorris ... Ruijia (Roger) Hu
-
Wyatt Dorris, et. al.Wyatt Dorris ... Ruijia (Roger) Hu
16 Mar 2020
16 Mar 2020

Detection of Hate Speech using BERT and Hate Speech Word Embedding with Deep Model
Hind Saleh ... Kawthar Moria
Applied Artificial Intelligence | VOL. 37
Hind Saleh, et. al.Hind Saleh ... Kawthar Moria
02 Feb 2023
Applied Artificial Intelligence | VOL. 37

Editage

Paperpal

R Discovery

Mind the Graph

R Discovery Prime

R Discovery Prime

Multilingual Hate Speech Detection: A Semi-Supervised Generative Adversarial Approach.

Abstract

Talk to us

Similar Papers

More From: Entropy