Read mapping on de Bruijn graphs.

Antoine Limasset,Eric Rivals,Bastien Cazaux,Pierre Peterlongo

doi:10.1186/s12859-016-1103-9

Abstract

BackgroundNext Generation Sequencing (NGS) has dramatically enhanced our ability to sequence genomes, but not to assemble them. In practice, many published genome sequences remain in the state of a large set of contigs. Each contig describes the sequence found along some path of the assembly graph, however, the set of contigs does not record all the sequence information contained in that graph. Although many subsequent analyses can be performed with the set of contigs, one may ask whether mapping reads on the contigs is as informative as mapping them on the paths of the assembly graph. Currently, one lacks practical tools to perform mapping on such graphs.ResultsHere, we propose a formal definition of mapping on a de Bruijn graph, analyse the problem complexity which turns out to be NP-complete, and provide a practical solution. We propose a pipeline called GGMAP (Greedy Graph MAPping). Its novelty is a procedure to map reads on branching paths of the graph, for which we designed a heuristic algorithm called BGREAT (de Bruijn Graph REAd mapping Tool). For the sake of efficiency, BGREAT rewrites a read sequence as a succession of unitigs sequences. GGMAP can map millions of reads per CPU hour on a de Bruijn graph built from a large set of human genomic reads. Surprisingly, results show that up to 22 % more reads can be mapped on the graph but not on the contig set.ConclusionsAlthough mapping reads on a de Bruijn graph is complex task, our proposal offers a practical solution combining efficiency with an improved mapping capacity compared to assembly-based mapping even for complex eukaryotic data.Electronic supplementary materialThe online version of this article (doi:10.1186/s12859-016-1103-9) contains supplementary material, which is available to authorized users.

Highlights

Generation Sequencing (NGS) has dramatically enhanced our ability to sequence genomes, but not to assemble them
We propose a more general problem, termed De Bruijn Graph Read Mapping Problem (DBGRMP), as we aim at mapping to a graph any source of Next Generation Sequencing (NGS) reads, either those reads used for building the graph or other reads
We introduce preliminary definitions, formalize the problem of mapping reads on paths of a de Bruijn graph (DBG), called the De Bruijn Graph Read Mapping Problem (DBGRMP), and prove it is NP-complete

Summary

Introduction

Generation Sequencing (NGS) has dramatically enhanced our ability to sequence genomes, but not to assemble them. Many subsequent analyses can be performed with the set of contigs, one may ask whether mapping reads on the contigs is as informative as mapping them on the paths of the assembly graph. Generation Sequencing technologies (NGS) have drastically accelerated the generation of sequenced genomes These technologies remain unable to provide a single sequence per chromosome. Instead, they produce a large and redundant set of reads, with each read being a piece of the whole genome. The assembly problem itself has been shown to be computationally difficult, more precisely NP-hard [2] Practical limitations arise both from the structure of Limasset et al BMC Bioinformatics (2016) 17:237 usually post-processed, for instance, by discarding short contigs

Objectives

Methods

Results

Talk to us

Join us for a 30 min session where you can share your feedback and ask us any queries you have

Schedule a call

Journal: BMC Bioinformatics	Publication Date: Jun 16, 2016
Citations: 90	License type: CC BY 4.0

R Discovery Prime

R Discovery Prime

Read mapping on de Bruijn graphs.

Abstract

Highlights

Summary

Talk to us

Similar Papers

More From: BMC Bioinformatics

Lead the way for us

Similar Papers

Efficient reconfiguration algorithms of de Bruijn and Kautz networks into linear arrays
Rabah Harbane ... Marie-Claude Heydemann
Theoretical Computer Science | VOL. 263
Rabah Harbane, et. al.Rabah Harbane ... Marie-Claude Heydemann
01 Jul 2001
Theoretical Computer Science | VOL. 263

De Bruijn Graph based De novo Genome Assembly
...
Journal of Software | VOL. -
, et. al. ...
08 Jan 2014
Journal of Software | VOL. -

An Optimized Graph-Based Metagenomic Gene Classification Approach
Md Sarwar Kamal ... Kaushik Dev
-
Md Sarwar Kamal, et. al.Md Sarwar Kamal ... Kaushik Dev
01 Jan 2020
01 Jan 2020

An Optimized Graph-Based Metagenomic Gene Classification Approach
Md Sarwar Kamal ... Linkon Chowdhury
-
Md Sarwar Kamal, et. al.Md Sarwar Kamal ... Linkon Chowdhury
01 Jan 2015
01 Jan 2015

Editage

Paperpal

R Discovery

Mind the Graph

R Discovery Prime

R Discovery Prime

Read mapping on de Bruijn graphs.

Abstract

Highlights

Summary

Talk to us

Similar Papers

More From: BMC Bioinformatics