Selecting dissimilar genes for multi-class classification, an application in cancer subtyping.

Zhipeng Cai,Mohammad R Salavatipour,Guohui Lin,Randy Goebel

doi:10.1186/1471-2105-8-206

Abstract

BackgroundGene expression microarray is a powerful technology for genetic profiling diseases and their associated treatments. Such a process involves a key step of biomarker identification, which are expected to be closely related to the disease. A most important task of these identified genes is that they can be used to construct a classifier which can effectively diagnose disease and even recognize the disease subtypes. Binary classification, for example, diseased or healthy, in microarray data analysis has been successful, while multi-class classification, such as cancer subtyping, remains challenging.ResultsWe target on the challenging multi-class classification in microarray data analysis, especially on the cancer subtyping using gene expression microarray. We present a novel class discrimination strength vector to represent individual genes and introduce a new measurement to quantify the class discrimination strength difference between two genes. Such a new distance measure is employed in gene clustering, and subsequently the gene cluster information is exploited to select a set of genes which can be used to construct a sample classifier.We tested our method on four real cancer microarray datasets each contains multiple subtypes of cancer patients. The experimental results show that the constructed classifiers all achieved a higher classification accuracy than the previously best classification results obtained on these four datasets. Additional tests show that the selected genes by our method are less correlated and they all contribute statistically significantly to the more accurate cancer subtyping.ConclusionThe proposed novel class discrimination strength vector is a better representation than the gene expression vector, in the sense that it can be used to effectively eliminate highly correlated but redundant genes for classifier construction. Such a method can build a classifier to achieve a higher classification accuracy, which is demonstrated via cancer subtyping.

Highlights

Gene expression microarray is a powerful technology for genetic profiling diseases and their associated treatments
The proposed novel class discrimination strength vector is a better representation than the gene expression vector, in the sense that it can be used to effectively eliminate highly correlated but redundant genes for classifier construction
Such a method can build a classifier to achieve a higher classification accuracy, which is demonstrated via cancer subtyping

Summary

Introduction

Gene expression microarray is a powerful technology for genetic profiling diseases and their associated treatments. Such a process involves a key step of biomarker identification, which are expected to be closely related to the disease. DNA microarray technology enables the measurement of expression levels of thousands of genes simultaneously This unique feature has a fundamental role in a wide range of current biological and medical research. In general, among the thousands of genes examined, only a small number of them are significantly associated with the experimental conditions These genes, or biomarkers, would have their expression levels increase or decrease under certain conditions, compared to normal expression levels. The huge number of genes versus a tiny number of experiments is a familiar machine learning challenge, often labeled as "the curse of dimensionality"

Methods

Results

Discussion

Conclusion

Full Text

Paper version not known

Open DOI Link

Talk to us

Join us for a 30 min session where you can share your feedback and ask us any queries you have

Schedule a call

Journal: BMC bioinformatics	Publication Date: Jun 16, 2007
Citations: 67	License type: cc-by

R Discovery Prime

R Discovery Prime

Selecting dissimilar genes for multi-class classification, an application in cancer subtyping.

Abstract

Highlights

Summary

Talk to us

Similar Papers

More From: BMC bioinformatics

Lead the way for us

Similar Papers

A comparative study of RNA-Seq and microarray data analysis on the two examples of rectal-cancer patients and Burkitt Lymphoma cells.
Alexander Wolff ... Petr V Nazarov
PloS one | VOL. 13
Alexander Wolff, et. al.Alexander Wolff ... Petr V Nazarov
16 May 2018
PloS one | VOL. 13

Microarrays and Epidemiology: Ensuring the Impact and Accessibility of Research Findings
Melissa A Troester ... Charles M Perou
Cancer Epidemiology, Biomarkers & Prevention | VOL. 18
Melissa A Troester, et. al.Melissa A Troester ... Charles M Perou
01 Jan 2009
Cancer Epidemiology, Biomarkers & Prevention | VOL. 18

3′-End Sequencing for Expression Quantification (3SEQ) from Archival Tumor Samples
Andrew H Beck ... Irene Oi Lin Ng
PLoS ONE | VOL. 5
Andrew H Beck, et. al.Andrew H Beck ... Irene Oi Lin Ng
19 Jan 2010
PLoS ONE | VOL. 5

Detection of COVID-19 cases through X-ray images using hybrid deep neural network
Rajit Nair ... Santosh Vishwakarma
World Journal of Engineering | VOL. 19
Rajit Nair, et. al.Rajit Nair ... Santosh Vishwakarma
11 Jan 2021
World Journal of Engineering | VOL. 19

Editage

Paperpal

R Discovery

Mind the Graph

R Discovery Prime

R Discovery Prime

Selecting dissimilar genes for multi-class classification, an application in cancer subtyping.

Abstract

Highlights

Summary

Talk to us

Similar Papers

More From: BMC bioinformatics