Genetic Programming for Preprocessing Tandem Mass Spectra to Improve the Reliability of Peptide Identification

Samaneh Azari,Mengjie Zhang,Lifeng Peng,Bing Xue

doi:10.1109/cec.2018.8477810

Abstract

Tandem mass spectrometry (MS/MS) is currently the most commonly used technology in proteomics for identifying proteins in complex biological samples. Mass spectrometers can produce a large number of MS/MS spectra each of which has hundreds of peaks. These peaks normally contain background noise, therefore a preprocessing step to filter the noise peaks can improve the accuracy and reliability of peptide identification. This paper proposes to preprocess the data by classifying peaks as noise peaks or signal peaks, i.e., a highly-imbalanced binary classification task, and uses genetic programming (GP) to address this task. The expectation is to increase the peptide identification reliability. Meanwhile, six different types of classification algorithms in addition to GP are used on various imbalance ratios and evaluated in terms of the average accuracy and recall. The GP method appears to be the best in the retention of more signal peaks as examined on a benchmark dataset containing 1, 674 MS/MS spectra. To further evaluate the effectiveness of the GP method, the preprocessed spectral data is submitted to a benchmark de novo sequencing software, PEAKS, to identify the peptides. The results show that the proposed method improves the reliability of peptide identification compared to the original un-preprocessed data and the intensity-based thresholding methods.

Talk to us

Join us for a 30 min session where you can share your feedback and ask us any queries you have

Schedule a call

R Discovery Prime

R Discovery Prime

Genetic Programming for Preprocessing Tandem Mass Spectra to Improve the Reliability of Peptide Identification

Abstract

Talk to us

Similar Papers

Lead the way for us

Similar Papers

Preprocessing Tandem Mass Spectra Using Genetic Programming for Peptide Identification.
Samaneh Azari ... Mengjie Zhang
Journal of the American Society for Mass Spectrometry | VOL. 30
Samaneh Azari, et. al.Samaneh Azari ... Mengjie Zhang
25 Apr 2019
Journal of the American Society for Mass Spectrometry | VOL. 30

مقایسه روش های برنامه ریزی ژنتیک و ماشین بردار پشتیبان در پیش بینی جریان روزانه رودخانه (مطالعه موردی: رودخانه باراندوزچای)
...
-
, et. al. ...
20 Oct 2014
20 Oct 2014

Enhanced Peptide Identification by Electron Transfer Dissociation Using an Improved Mascot Percolator
James C Wright ... Jyoti S Choudhary
Molecular & Cellular Proteomics | VOL. 11
James C Wright, et. al.James C Wright ... Jyoti S Choudhary
01 Aug 2012
Molecular & Cellular Proteomics | VOL. 11

Peptizer, a Tool for Assessing False Positive Peptide Identifications and Manually Validating Selected Results
Kenny Helsens ... Lennart Martens
Molecular & Cellular Proteomics | VOL. 7
Kenny Helsens, et. al.Kenny Helsens ... Lennart Martens
01 Dec 2008
Molecular & Cellular Proteomics | VOL. 7

Editage

Paperpal

R Discovery

Mind the Graph

R Discovery Prime

R Discovery Prime

Genetic Programming for Preprocessing Tandem Mass Spectra to Improve the Reliability of Peptide Identification

Abstract

Talk to us

Similar Papers