Improving data splitting for classification applications in spectrochemical analyses employing a random-mutation Kennard-Stone algorithm approach.

Camilo L M Morais,Marfran C D Santos,Francis L Martin,Kássio M G Lima,Jonathan Wren

doi:10.1093/bioinformatics/btz421

Camilo L M Morais, Marfran C D Santos + Show 3 more

Open Access

https://doi.org/10.1093/bioinformatics/btz421

Copy DOI

Abstract

MotivationData splitting is a fundamental step for building classification models with spectral data, especially in biomedical applications. This approach is performed following pre-processing and prior to model construction, and consists of dividing the samples into at least training and test sets; herein, the training set is used for model construction and the test set for model validation. Some of the most-used methodologies for data splitting are the random selection (RS) and the Kennard-Stone (KS) algorithms; here, the former works based on a random splitting process and the latter is based on the calculation of the Euclidian distance between the samples. We propose an algorithm called the Morais-Lima-Martin (MLM) algorithm, as an alternative method to improve data splitting in classification models. MLM is a modification of KS algorithm by adding a random-mutation factor.ResultsRS, KS and MLM performance are compared in simulated and six real-world biospectroscopic applications using principal component analysis linear discriminant analysis (PCA-LDA). MLM generated a better predictive performance in comparison with RS and KS algorithms, in particular regarding sensitivity and specificity values. Classification is found to be more well-equilibrated using MLM. RS showed the poorest predictive response, followed by KS which showed good accuracy towards prediction, but relatively unbalanced sensitivities and specificities. These findings demonstrate the potential of this new MLM algorithm as a sample selection method for classification applications in comparison with other regular methods often applied in this type of data.Availability and implementationMLM algorithm is freely available for MATLAB at https://doi.org/10.6084/m9.figshare.7393517.v1.

Highlights

Data splitting is a process used to separate a given dataset into at least two subsets called ‘training’ and ‘test’
This is made by using chemometric methods such as principal component analysis linear discriminant analysis (PCA-LDA) (Morais and Lima, 2018), partial least squares discriminant analysis (PLS-DA) (Brereton and Lloyd, 2014), or support vector machines (SVM) (Cortes and Vapnik, 1995)
This is made by firstly removing a certain number of samples from the training set and building the classification model with the remaining samples, where the removed samples are predicted as a temporary validation set

Summary

Introduction

Data splitting is a process used to separate a given dataset into at least two subsets called ‘training’ (or ‘calibration’) and ‘test’ (or ‘prediction’). MLM generated a better predictive performance in comparison with RS and KS algorithms, in particular regarding sensitivity and specificity values. These findings demonstrate the potential of this new MLM algorithm as a sample selection method for classification applications in comparison with other regular methods often applied in this type of data.

Results

Conclusion

Full Text

Paper version not known

Open DOI Link

Talk to us

Join us for a 30 min session where you can share your feedback and ask us any queries you have

Schedule a call

Journal: Bioinformatics	Publication Date: May 22, 2019
Citations: 76	License type: CC BY 4.0

R Discovery Prime

R Discovery Prime

Improving data splitting for classification applications in spectrochemical analyses employing a random-mutation Kennard-Stone algorithm approach.

Abstract

Highlights

Summary

Talk to us

Similar Papers

More From: Bioinformatics

Lead the way for us

Similar Papers

Prediction of placenta barrier permeability and reproductive toxicity of compounds in tocolytic Chinese herbs using support vector machine
Fang Lu ... Yusu He
-
Fang Lu, et. al.Fang Lu ... Yusu He
01 Jan 2015
01 Jan 2015

A comparative study of representative subset selection for NIR model updating
Xiaping Fu ... Yibin Ying
-
Xiaping Fu, et. al. Xiaping Fu ... Yibin Ying
01 Jan 2010
01 Jan 2010

Impact assessment of the rational selection of training and test sets on the predictive ability of QSAR models
M F Andrada ... J C Garro Martinez
SAR and QSAR in Environmental Research | VOL. 28
M F Andrada, et. al.M F Andrada ... J C Garro Martinez
14 Nov 2017
SAR and QSAR in Environmental Research | VOL. 28

P.1.c.010 Gender-dependent influence of bupropion pretreatment on locomotor activity and the activity of hepatic CYP2D2 in rats
L Zahradníková ... E Hadaova
European Neuropsychopharmacology | VOL. 17
L Zahradníková, et. al.L Zahradníková ... E Hadaova
01 Oct 2007
European Neuropsychopharmacology | VOL. 17

Editage

Paperpal

R Discovery

Mind the Graph

R Discovery Prime

R Discovery Prime

Improving data splitting for classification applications in spectrochemical analyses employing a random-mutation Kennard-Stone algorithm approach.

Abstract

Highlights

Summary

Talk to us

Similar Papers

More From: Bioinformatics