Comparative analysis of missing value imputation methods to improve clustering and interpretation of microarray experiments

Magalie Celton,Gaëlle Lelandais,Alain Malpertuy,Alexandre G De Brevern

doi:10.1186/1471-2164-11-15

Abstract

BackgroundMicroarray technologies produced large amount of data. In a previous study, we have shown the interest of k-Nearest Neighbour approach for restoring the missing gene expression values, and its positive impact of the gene clustering by hierarchical algorithm. Since, numerous replacement methods have been proposed to impute missing values (MVs) for microarray data. In this study, we have evaluated twelve different usable methods, and their influence on the quality of gene clustering. Interestingly we have used several datasets, both kinetic and non kinetic experiments from yeast and human.ResultsWe underline the excellent efficiency of approaches proposed and implemented by Bo and co-workers and especially one based on expected maximization (EM_array). These improvements have been observed also on the imputation of extreme values, the most difficult predictable values. We showed that the imputed MVs have still important effects on the stability of the gene clusters. The improvement on the clustering obtained by hierarchical clustering remains limited and, not sufficient to restore completely the correct gene associations. However, a common tendency can be found between the quality of the imputation method and the gene cluster stability. Even if the comparison between clustering algorithms is a complex task, we observed that k-means approach is more efficient to conserve gene associations.ConclusionsMore than 6.000.000 independent simulations have assessed the quality of 12 imputation methods on five very different biological datasets. Important improvements have so been done since our last study. The EM_array approach constitutes one efficient method for restoring the missing expression gene values, with a lower estimation error level. Nonetheless, the presence of MVs even at a low rate is a major factor of gene cluster instability. Our study highlights the need for a systematic assessment of imputation methods and so of dedicated benchmarks. A noticeable point is the specific influence of some biological dataset.

Highlights

Microarray technologies produced large amount of data
Simulated missing values are generated for a fixed τ percentage and are included in the Reference matrix
Microarrays studies must take into account the important problem of missing values for the validity of biological results

Summary

Introduction

Microarray technologies produced large amount of data. In a previous study, we have shown the interest of k-Nearest Neighbour approach for restoring the missing gene expression values, and its positive impact of the gene clustering by hierarchical algorithm. Numerous replacement methods have been proposed to impute missing values (MVs) for microarray data. We have evaluated twelve different usable methods, and their influence on the quality of gene clustering. During the image analysis phase, corrupted or suspicious spots are filtered [11], generating missing data [18]. These missing values (MVs) disturb the gene clustering obtained by classical clustering methods, e.g., hierarchical clustering [19], k-means clustering [20], Kohonen Maps [21,22] or projection methods, e.g., Principal Component Analysis [23]. To limit skews related to the MVs, several methodologies using the values present in the data file to replace the MVs by estimated values have been developed [25]

Objectives

Results

Discussion

Conclusion

Talk to us

Join us for a 30 min session where you can share your feedback and ask us any queries you have

Schedule a call

Journal: BMC Genomics	Publication Date: Jan 7, 2010
Citations: 167	License type: cc-by

R Discovery Prime

R Discovery Prime

Comparative analysis of missing value imputation methods to improve clustering and interpretation of microarray experiments

Abstract

Highlights

Summary

Talk to us

Similar Papers

More From: BMC Genomics

Lead the way for us

Similar Papers

A New Paradigm for Development of Data Imputation Approach for Missing Value Estimation
Madhu G ... Nagachandrika G
-
Madhu G, et. al.Madhu G ... Nagachandrika G
01 Dec 2016
01 Dec 2016

A New Paradigm for Development of Data Imputation Approach for Missing Value Estimation
Madhu G ... Nagachandrika G
International Journal of Electrical and Computer Engineering (IJECE) | VOL. 6
Madhu G, et. al.Madhu G ... Nagachandrika G
01 Dec 2016
International Journal of Electrical and Computer Engineering (IJECE) | VOL. 6

Influence of microarrays experiments missing values on the stability of gene groups by hierarchical clustering
Alexandre G De Brevern ... Alain Malpertuy
BMC Bioinformatics | VOL. 5
Alexandre G De Brevern, et. al.Alexandre G De Brevern ... Alain Malpertuy
01 Jan 2004
BMC Bioinformatics | VOL. 5

Comparative Study of Four Methods in Missing Value Imputations under Missing Completely at Random Mechanism
Michikazu Nakai ... Yoshihiro Miyamoto
Open Journal of Statistics | VOL. 04
Michikazu Nakai, et. al.Michikazu Nakai ... Yoshihiro Miyamoto
01 Jan 2014
Open Journal of Statistics | VOL. 04

Editage

Paperpal

R Discovery

Mind the Graph

R Discovery Prime

R Discovery Prime

Comparative analysis of missing value imputation methods to improve clustering and interpretation of microarray experiments

Abstract

Highlights

Summary

Talk to us

Similar Papers

More From: BMC Genomics