Multiple organism algorithm for finding ultraconserved elements

Scott Christley,Neil F Lobo,Greg Madey

doi:10.1186/1471-2105-9-15

Scott Christley, Neil F Lobo + Show 1 more

Open Access

https://doi.org/10.1186/1471-2105-9-15

Copy DOI

Journal: BMC Bioinformatics	Publication Date: Jan 11, 2008
Citations: 40	License type: CC BY 2.0

Affiliation: University of Notre Dame, Biocom

Abstract

BackgroundUltraconserved elements are nucleotide or protein sequences with 100% identity (no mismatches, insertions, or deletions) in the same organism or between two or more organisms. Studies indicate that these conserved regions are associated with micro RNAs, mRNA processing, development and transcription regulation. The identification and characterization of these elements among genomes is necessary for the further understanding of their functionality.ResultsWe describe an algorithm and provide freely available software which can find all of the ultraconserved sequences between genomes of multiple organisms. Our algorithm takes a combinatorial approach that finds all sequences without requiring the genomes to be aligned. The algorithm is significantly faster than BLAST and is designed to handle very large genomes efficiently. We ran our algorithm on several large comparative analyses to evaluate its effectiveness; one compared 17 vertebrate genomes where we find 123 ultraconserved elements longer than 40 bps shared by all of the organisms, and another compared the human body louse, Pediculus humanus humanus, against itself and select insects to find thousands of non-coding, potentially functional sequences.ConclusionWhole genome comparative analysis for multiple organisms is both feasible and desirable in our search for biological knowledge. We argue that bioinformatic programs should be forward thinking by assuming analysis on multiple (and possibly large) genomes in the design and implementation of algorithms. Our algorithm shows how a compromise design with a trade-off of disk space versus memory space allows for efficient computation while only requiring modest computer resources, and at the same time providing benefits not available with other software.

Highlights

Ultraconserved elements are nucleotide or protein sequences with 100% identity in the same organism or between two or more organisms
The main algorithm takes two suffix arrays and produces a list of maximal common prefixes (MCP) which correspond to ultraconserved sequences; this is done in a pair-wise fashion for all suffix arrays of the two organisms
Each final MCP file for a pair of organisms is intersected within another final MCP file for a different pair of organisms; this is repeated in tournament-style fashion producing an MCP file for all the organisms

Summary

Methodology article

Address: 1Department of Computer Science and Engineering, University of Notre Dame, Notre Dame, IN, 46556, USA, 2Interdisciplinary Center for the Study of Biocomplexity, University of Notre Dame, USA and 3Department of Biological Sciences, University of Notre Dame, Notre Dame, IN, 46556, USA. Published: 11 January 2008 BMC Bioinformatics 2008, 9:15 doi:10.1186/1471-2105-9-15

Results

Conclusion

Background

Results and Discussion

18. Kent WJ

28. Gusfield D

Full Text

Paper version not known

Open DOI Link

Talk to us

Join us for a 30 min session where you can share your feedback and ask us any queries you have

Schedule a call

R Discovery Prime

R Discovery Prime

Multiple organism algorithm for finding ultraconserved elements

Abstract

Highlights

Summary

Talk to us

Similar Papers

More From: BMC Bioinformatics

Lead the way for us

Similar Papers

Combinatorial Gene Regulatory Functions Underlie Ultraconserved Elements in Drosophila.
Maria Warnefors ... Claudio R Alonso
Molecular Biology and Evolution | VOL. 33
Maria Warnefors, et. al.Maria Warnefors ... Claudio R Alonso
31 May 2016
Molecular Biology and Evolution | VOL. 33

Ultraconserved Elements ( UCEs ) in the Human Genome
Alison P Lee ... B Venkatesh
-
Alison P Lee, et. al.Alison P Lee ... B Venkatesh
15 Apr 2013
15 Apr 2013

Large-Scale Appearance of Ultraconserved Elements in Tetrapod Genomes and Slowdown of the Molecular Clock
S Stephen ... I V Makunin
Molecular Biology and Evolution | VOL. 25
S Stephen, et. al.S Stephen ... I V Makunin
02 Jan 2008
Molecular Biology and Evolution | VOL. 25

Evolutionary growth process of highly conserved sequences in vertebrate genomes
Minaka Ishibashi ... Tadashi Imanishi
Gene | VOL. 504
Minaka Ishibashi, et. al.Minaka Ishibashi ... Tadashi Imanishi
09 May 2012
Gene | VOL. 504

Editage

Paperpal

R Discovery

Mind the Graph

R Discovery Prime

R Discovery Prime

Multiple organism algorithm for finding ultraconserved elements

Abstract

Highlights

Summary

Talk to us

Similar Papers

More From: BMC Bioinformatics