A Strategy for Training Set Selection in Text Classification Problems

Maria Luiza,Nelson F,Katiusca B,Grazziela P

doi:10.14569/ijacsa.2013.040608

Abstract

An issue in text classification problems involves the choice of good samples on which to train the classifier. Training sets that properly represent the characteristics of each class have a better chance of establishing a successful predictor. Moreover, sometimes data are redundant or take large amounts of computing time for the learning process. To overcome this issue, data selection techniques have been proposed, including instance selection. Some data mining techniques are based on nearest neighbors, ordered removals, random sampling, particle swarms or evolutionary methods. The weaknesses of these methods usually involve a lack of accuracy, lack of robustness when the amount of data increases, over?tting and a high complexity. This work proposes a new immune-inspired suppressive mechanism that involves selection. As a result, data that are not relevant for a classifier’s ?nal model are eliminated from the training process. Experiments show the e?ectiveness of this method, and the results are compared to other techniques; these results show that the proposed method has the advantage of being accurate and robust for large data sets, with less complexity in the algorithm.

Highlights

Nowadays most of the information is stored electronically, in the form of text databases
This paper proposes a new approach for addressing the training data reduction in text mining classifications problems
The performance of the two classification algorithms Naive Bayes and Support Vector Machine (SVM) over the resulting reduced training and test subsets of SeleSup is compared to the performance over the subsets selected by the CHC algorithm, which is based on genetic algorithms [19] and random sampling (RS) based on the reduction percentages of experiments of each algorithm

Summary

A Strategy for Training Set Selection in Text Classification Problems

Sometimes data are redundant or take large amounts of computing time for the learning process. To overcome this issue, data selection techniques have been proposed, including instance selection. Some data mining techniques are based on nearest neighbors, ordered removals, random sampling, particle swarms or evolutionary methods. The weaknesses of these methods usually involve a lack of accuracy, lack of robustness when the amount of data increases, overfitting and a high complexity. Experiments show the effectiveness of this method, and the results are compared to other techniques; these results show that the proposed method has the advantage of being accurate and robust for large data sets, with less complexity in the algorithm

INTRODUCTION

OBJECTIVES

PREVIOUS WORK

THE SUPPRESSION MECHANISM

EXPERIMENTAL STUDY

REUTERS

Newsgroup Data

Parameters

Significance Test

RESULTS AND ANALYSIS

RESULTS

VIII. CONCLUSION

FUTURE WORK

Full Text

Paper version not known

Open DOI Link

Talk to us

Join us for a 30 min session where you can share your feedback and ask us any queries you have

Schedule a call

Journal: International Journal of Advanced Computer Science and Applications	Publication Date: Jan 1, 2013
Citations: 27	License type: cc-by

R Discovery Prime

R Discovery Prime

A Strategy for Training Set Selection in Text Classification Problems

Abstract

Highlights

Summary

Talk to us

Similar Papers

More From: International Journal of Advanced Computer Science and Applications

Lead the way for us

Similar Papers

An immune-inspired instance selection mechanism for supervised classification
Grazziela P Figueredo ... Nelson F F Ebecken
Memetic Computing | VOL. 4
Grazziela P Figueredo, et. al.Grazziela P Figueredo ... Nelson F F Ebecken
31 Mar 2012
Memetic Computing | VOL. 4

Improved particle swarm optimization algorithm using design of experiment and data mining techniques
Zhao Liu ... Ren-Jye Yang
Structural and Multidisciplinary Optimization | VOL. 52
Zhao Liu, et. al.Zhao Liu ... Ren-Jye Yang
01 Jul 2015
Structural and Multidisciplinary Optimization | VOL. 52

Prototype Generation Using Multiobjective Particle Swarm Optimization for Nearest Neighbor Classification.
Weiwei Hu ... Ying Tan
IEEE Transactions on Cybernetics | VOL. 46
Weiwei Hu, et. al.Weiwei Hu ... Ying Tan
19 Oct 2015
IEEE Transactions on Cybernetics | VOL. 46

Trainable segmentation for transmission electron microscope images of inorganic nanoparticles.
Cameron G Bell ... Manfred E Schuster
Journal of Microscopy | VOL. 288
Cameron G Bell, et. al.Cameron G Bell ... Manfred E Schuster
11 May 2022
Journal of Microscopy | VOL. 288

Editage

Paperpal

R Discovery

Mind the Graph

R Discovery Prime

R Discovery Prime

A Strategy for Training Set Selection in Text Classification Problems

Abstract

Highlights

Summary

Talk to us

Similar Papers

More From: International Journal of Advanced Computer Science and Applications