A Comparative Study of Content Statistics of Coding Regions in an Evolutionary Computation Framework for Gene Prediction

Javier Pérez-Rodríguez,Alexis G Arroyo-Peña,Nicolás García-Pedrajas

doi:10.1007/978-3-642-31087-4_22

Abstract

AbstractThe determination of which parts of a DNA sequence are coding is an unsolved and relevant problem in the field of bioinformatics. This problem is called gene prediction or gene finding, and it consists of locating the most likely gene structure in a genomic sequence.Taking into account some restrictions, gene structure prediction may be considered as a search problem. To address the problem, evolutionary computation approaches can be used, although their performance will depend on the discriminative power of the statistical measures employed to extract useful features from the sequence.In this study, we test six different content statistics to determine which of them have higher relevance in an evolutionary search for coding and non-coding regions of human DNA. We conduct this comparative study on the human chromosomes 3, 19 and 21.KeywordsCodon UsageSynonymous CodonContent StatisticAverage Mutual InformationTranslation Initiation SiteThese keywords were added by machine and not by the authors. This process is experimental and the keywords may be updated as the learning algorithm improves.

Full Text