Compressing population DNA sequences using multiple reference sequences

Kin-On Cheng,Ngai-Fong Law,Wan-Chi Siu

doi:10.1109/apsipa.2017.8282136

Abstract

Compressing population DNA sequences often relies on the use of a reference sequence so that only the differences between the target DNA sequences to be compressed and the reference sequence are encoded. Despite the importance of the choice of the reference sequence, state-of-the-art algorithms in population sequence compression often selected one of the population sequences as a reference sequence in an ad hoc manner. In this paper, we investigated issues about the choice of the reference sequence. In particular, population sequences are first clustered into a number of groups. A reference sequence is then obtained for each group so that substructures within each group can be characterized by this reference sequence. Afterwards, the reference sequence is used to compress sequences within that group. In this way, the multiple reference sequences framework can optimize the overall compression performance on the set of population sequences. Results show that our proposed method reduces the compressed size by up to 91% as compared to state-of-the-art reference- based approaches.

Talk to us

Join us for a 30 min session where you can share your feedback and ask us any queries you have

Schedule a call

R Discovery Prime

R Discovery Prime

Compressing population DNA sequences using multiple reference sequences

Abstract

Talk to us

Similar Papers

Lead the way for us

Similar Papers

Compression of Multiple DNA Sequences Using Intra-Sequence and Inter-Sequence Similarities.
Kin-On Cheng ... Ngai-Fong Law
IEEE/ACM transactions on computational biology and bioinformatics | VOL. 12
Kin-On Cheng, et. al.Kin-On Cheng ... Ngai-Fong Law
01 Nov 2015
IEEE/ACM transactions on computational biology and bioinformatics | VOL. 12

Author response: A large gene family in fission yeast encodes spore killers that subvert Mendel’s law
Wen Hu ... Li-Lin Du
-
Wen Hu, et. al.Wen Hu ... Li-Lin Du
02 May 2017
02 May 2017

Differential Gene Expression in the Siphonophore Nanomia bijuga (Cnidaria) Assessed with Multiple Next-Generation Sequencing Workflows
Stefan Siebert ... Sophia C Tintori
PLoS ONE | VOL. 6
Stefan Siebert, et. al.Stefan Siebert ... Sophia C Tintori
29 Jul 2011
PLoS ONE | VOL. 6

Integration of Alignment and Phylogeny in the Whole-Genome Era

-

18 Jun 2015
18 Jun 2015

Editage

Paperpal

R Discovery

Mind the Graph

R Discovery Prime

R Discovery Prime

Compressing population DNA sequences using multiple reference sequences

Abstract

Talk to us

Similar Papers