Non-overlapping Common Substrings Allowing Mutations

H L Chan,S M Yiu,W K Sung,T W Lam,P W H Wong

doi:10.1007/s11786-007-0030-6

Abstract

This paper studies several combinatorial problems arising from finding the conserved genes of two genomes (i.e., the entire DNA of two species). The input is a collection of n maximal common substrings of the two genomes. The problem is to find, based on different criteria, a subset of such common substrings with maximum total length. The most basic criterion requires that the common substrings selected have the same ordering in the two genomes and they do not overlap among themselves in either genome. To capture mutations (transpositions and reversals) between the genomes, we do not insist the substrings selected to have the same ordering. Conceptually, we allow one ordering to go through some mutations to become the other ordering. If arbitrary mutations are allowed, the problem of finding a maximum-length, non-overlapping subset of substrings is found to be NP-hard. However, arbitrary mutations probably overmodel the problem and are likely to find more noise than conserved genes. We consider two criteria that attempt to model sparse and non-overlapping mutations. We show that both can be solved in polynomial time using dynamic programming.

Talk to us

Join us for a 30 min session where you can share your feedback and ask us any queries you have

Schedule a call

R Discovery Prime

R Discovery Prime

Non-overlapping Common Substrings Allowing Mutations

Abstract

Talk to us

Similar Papers

More From: Mathematics in Computer Science

Lead the way for us

Journal: Mathematics in Computer Science	Publication Date: Apr 1, 2008
Citations: 14

Similar Papers

Variations in biological characteristics of temperate gonochoristic species of Platycephalidae and their implications: A review
Peter G Coulson ... Ian C Potter
Estuarine, Coastal and Shelf Science | VOL. 190
Peter G Coulson, et. al.Peter G Coulson ... Ian C Potter
27 Mar 2017
Estuarine, Coastal and Shelf Science | VOL. 190

Longest Common Substring with Approximately k Mismatches

-

01 Jun 2016
01 Jun 2016

Tadpole of Telmatobius mayoloi (Anura: Ceratophryidae)
César Aguilar ... Edgar Lehr
Journal of Herpetology | VOL. 43
César Aguilar, et. al.César Aguilar ... Edgar Lehr
01 Mar 2009
Journal of Herpetology | VOL. 43

On the Common Substring Alignment Problem
Gad M Landau ... Michal Ziv-Ukelson
Journal of Algorithms | VOL. 41
Gad M Landau, et. al.Gad M Landau ... Michal Ziv-Ukelson
01 Nov 2001
Journal of Algorithms | VOL. 41

Editage

Paperpal

R Discovery

Mind the Graph

R Discovery Prime

R Discovery Prime

Non-overlapping Common Substrings Allowing Mutations

Abstract

Talk to us

Similar Papers

More From: Mathematics in Computer Science