Abstract

Feature selection on mass spectrometry (MS) data is essential for improving classification performance and biomarker discovery. The number of MS samples is typically very small compared with the high dimensionality of the samples, which makes the problem of biomarker discovery very hard. In this paper, we propose the use of genetic programming for biomarker detection and classification of MS data. The proposed approach is composed of two phases: in the first phase, feature selection and ranking are performed. In the second phase, classification is performed. The results show that the proposed method can achieve better classification performance and biomarker detection rate than the information gain- (IG) based and the RELIEF feature selection methods. Meanwhile, four classifiers, Naive Bayes, J48 decision tree, random forest and support vector machines, are also used to further test the performance of the top ranked features. The results show that the four classifiers using the top ranked features from the proposed method achieve better performance than the IG and the RELIEF methods. Furthermore, GP also outperforms a genetic algorithm approach on most of the used data sets.

Full Text
Published version (Free)

Talk to us

Join us for a 30 min session where you can share your feedback and ask us any queries you have

Schedule a call