A Toxic Comment Classification Model Based on Ensemble

Jian Xu,Yuqing Zhai

doi:10.1088/1742-6596/1873/1/012080

Abstract

Accurate classification of toxic comments in different languages is an important task in today’s international social networking platforms. In order to improve the classification of toxic comments in different languages, this paper combines the advantages of high accuracy of monolingual model and strong generalization ability of multi-language model, and adopts the ensemble of multilingual model and monolingual model to classify toxic comments. For the monolingual model, the monolingual pre-training model is fine-tuned with labeled task data; for the multilingual model, before fine-tuning the model with labeled task data, a further pre-training is applied to train model on unlabeled data, which aims to make full use of the large amount of unlabeled data and reduce the dependency of amount of labeled comment data while improving the classification effect. Comparative experiments on Conversasion AI’s multilingual toxic comment dataset show that the model in this paper has improved results on different evaluation metrics compared to the XLM-RoBERTa multilingual fine-tuning model, illustrating the validity of the model.

Talk to us

Join us for a 30 min session where you can share your feedback and ask us any queries you have

Schedule a call

Journal: Journal of Physics: Conference Series	Publication Date: Apr 1, 2021
Citations: 1	License type: cc-by

R Discovery Prime

R Discovery Prime

A Toxic Comment Classification Model Based on Ensemble

Abstract

Talk to us

Similar Papers

More From: Journal of Physics: Conference Series

Lead the way for us

Similar Papers

Improving sentence representation for vietnamese natural language understanding using optimal transport
Phu Xuan-Vinh Nguyen ... Kiet Van Nguyen
Journal of Intelligent & Fuzzy Systems | VOL. -
Phu Xuan-Vinh Nguyen, et. al.Phu Xuan-Vinh Nguyen ... Kiet Van Nguyen
27 Jun 2023
Journal of Intelligent & Fuzzy Systems | VOL. -

NASca and NASes: Two Monolingual Pre-Trained Models for Abstractive Summarization in Catalan and Spanish
Vicent Ahuir ... Encarna Segarra
Applied Sciences | VOL. 11
Vicent Ahuir, et. al.Vicent Ahuir ... Encarna Segarra
22 Oct 2021
Applied Sciences | VOL. 11

Dataset Augmentation for Counteracting Bias in Toxic Comment Classification
Senhao Cheng
Highlights in Science, Engineering and Technology | VOL. 85
Senhao ChengSenhao Cheng
13 Mar 2024
Highlights in Science, Engineering and Technology | VOL. 85

IMPACT OF THE SYNTACTIC DEPENDENCIES IN THE SENTENCES ON THE QUALITY OF THE IDENTIFICATION OF THE TOXIC COMMENTS IN THE SOCIAL NETWORKS
S D Shtovba ... М V Petrychko
Scientific Works of Vinnytsia National Technical University | VOL. -
S D Shtovba, et. al.S D Shtovba ... М V Petrychko
01 Jan 2019
Scientific Works of Vinnytsia National Technical University | VOL. -

Editage

Paperpal

R Discovery

Mind the Graph

R Discovery Prime

R Discovery Prime

A Toxic Comment Classification Model Based on Ensemble

Abstract

Talk to us

Similar Papers

More From: Journal of Physics: Conference Series