Forward feature selection for toxic speech classification using support vector machine and random forest

Agustinus Bimo Gumelar,Derry Pramono Adi,Indar Sugiarto,Frismanda Frismanda,Astri Yogatama

doi:10.11591/ijai.v11.i2.pp717-726

Agustinus Bimo Gumelar, Derry Pramono Adi + Show 3 more

Open Access

https://doi.org/10.11591/ijai.v11.i2.pp717-726

Copy DOI

Abstract

<span lang="EN-US">This study describes the methods for eliminating irrelevant features in speech data to enhance toxic speech classification accuracy and reduce the complexity of the learning process. Therefore, the wrapper method is introduced to estimate the forward selection technique based on support vector machine (SVM) and random forest (RF) classifier algorithms. Eight main speech features were then extracted with derivatives consisting of 9 statistical sub-features from 72 features in the extraction process. Furthermore, Python is used to implement the classifier algorithm of 2,000 toxic data collected through the world's largest video sharing media, known as YouTube. Conclusively, this experiment shows that after the feature selection process, the classification performance using SVM and RF algorithms increases to an excellent extent. We were able to select 10 speech features out of 72 original feature sets using the forward feature selection method, with 99.5% classification accuracy using RF and 99.2% using SVM.</span>

Talk to us

Join us for a 30 min session where you can share your feedback and ask us any queries you have

Schedule a call

Journal: IAES International Journal of Artificial Intelligence (IJ-AI)	Publication Date: Jun 1, 2022
Citations: 2	License type: CC BY-SA 4.0

R Discovery Prime

R Discovery Prime

Forward feature selection for toxic speech classification using support vector machine and random forest

Abstract

Talk to us

Similar Papers

More From: IAES International Journal of Artificial Intelligence (IJ-AI)

Lead the way for us

Similar Papers

Methods of forward feature selection based on the aggregation of classifiers generated by single attribute
Linkai Luo ... Fan Yang
Computers in Biology and Medicine | VOL. 41
Linkai Luo, et. al.Linkai Luo ... Fan Yang
07 May 2011
Computers in Biology and Medicine | VOL. 41

Object based classification of benthic habitat using Sentinel 2 imagery by applying with support vector machine and random forest algorithms in shallow waters of Kepulauan Seribu, Indonesia
Hartoni Hartoni ... Syamsul Bahri Agus
Biodiversitas Journal of Biological Diversity | VOL. 23
Hartoni Hartoni, et. al.Hartoni Hartoni ... Syamsul Bahri Agus
07 Jan 2022
Biodiversitas Journal of Biological Diversity | VOL. 23

A machine learning framework for spatio-temporal vulnerability mapping of groundwaters to nitrate in a data scarce region in Lenjanat Plain, Iran.
Reza Jalali ... Hossein Hashemi
Environmental science and pollution research international | VOL. 31
Reza Jalali, et. al.Reza Jalali ... Hossein Hashemi
11 Jun 2024
Environmental science and pollution research international | VOL. 31

Use of machine learning-based classification algorithms in the monitoring of Land Use and Land Cover practices in a hilly terrain.
Deepanshu Parashar ... Ajit Pratap Singh
Environmental Monitoring and Assessment | VOL. 196
Deepanshu Parashar, et. al.Deepanshu Parashar ... Ajit Pratap Singh
05 Dec 2023
Environmental Monitoring and Assessment | VOL. 196

Editage

Paperpal

R Discovery

Mind the Graph

R Discovery Prime

R Discovery Prime

Forward feature selection for toxic speech classification using support vector machine and random forest

Abstract

Talk to us

Similar Papers

More From: IAES International Journal of Artificial Intelligence (IJ-AI)