Clustering Gene-Expression Data: A Hybrid Approach that Iterates Between k-Means and Evolutionary Search

E R Hruschka,R J G B Campello,L N De Castro

doi:10.1007/978-3-540-73297-6_12

Abstract

Summary. Clustering genes based on their expression profiles is usually the first step in geneexpression data analysis. Among the many algorithms that can be applied to gene clustering, the k-means algorithm is one of the most popular techniques. This is mainly due to its ease of comprehension, implementation, and interpretation of the results. However, k-means suffers from some problems, such as the need to define a priori the number of clusters (k )a nd the possibility of getting trapped into local optimal solutions. Evolutionary algorithms for clustering, by contrast, are known for being capable of performing broad searches over the space of possible solutions and can be used to automatically estimate the number of clusters. This work elaborates on an evolutionary algorithm specially designed to solve clustering problems and shows how it can be used to optimize the k-means algorithm. The performance of the resultant hybrid approach is illustrated by means of experiments in several bioinformatics datasets with multiple measurements, which are expected to yield more accurate and more stable clusters. Two different measures (Euclidean and Pearson) are employed for computing (dis)similarities between genes. A review of the use of evolutionary algorithms for gene-expression data processing is also included.

Talk to us

Join us for a 30 min session where you can share your feedback and ask us any queries you have

Schedule a call

R Discovery Prime

R Discovery Prime

Clustering Gene-Expression Data: A Hybrid Approach that Iterates Between k-Means and Evolutionary Search

Abstract

Talk to us

Similar Papers

Lead the way for us

Similar Papers

Evolving clusters in gene-expression data
Eduardo R Hruschka ... Leandro N De Castro
Information Sciences | VOL. 176
Eduardo R Hruschka, et. al.Eduardo R Hruschka ... Leandro N De Castro
29 Aug 2005
Information Sciences | VOL. 176

Comparative Analysis of Different Label-Free Mass Spectrometry Based Protein Abundance Estimates and Their Correlation with RNA-Seq Gene Expression Data
Kang Ning ... Damian Fermin
Journal of Proteome Research | VOL. 11
Kang Ning, et. al.Kang Ning ... Damian Fermin
29 Feb 2012
Journal of Proteome Research | VOL. 11

Evolutionary Fuzzy Clustering: An Overview and Efficiency Issues
D Horta ... E R Hruschka
-
D Horta, et. al.D Horta ... E R Hruschka
01 Jan 2009
01 Jan 2009

Multiobjective clustering with automatic k-determination for large-scale data
Nobukazu Matake ... Tomoyuki Hiroyasu
-
Nobukazu Matake, et. al.Nobukazu Matake ... Tomoyuki Hiroyasu
07 Jul 2007
07 Jul 2007

Editage

Paperpal

R Discovery

Mind the Graph

R Discovery Prime

R Discovery Prime

Clustering Gene-Expression Data: A Hybrid Approach that Iterates Between k-Means and Evolutionary Search

Abstract

Talk to us

Similar Papers