Assessment of k-mer spectrum applicability for metagenomic dissimilarity analysis.

Veronika B Dubinkina,Dmitry S Ischenko,Alexander V Tyakht,Dmitry G Alexeev,Vladimir I Ulyantsev

doi:10.1186/s12859-015-0875-7

Abstract

BackgroundA rapidly increasing flow of genomic data requires the development of efficient methods for obtaining its compact representation. Feature extraction facilitates classification, clustering and model analysis for testing and refining biological hypotheses. “Shotgun” metagenome is an analytically challenging type of genomic data - containing sequences of all genes from the totality of a complex microbial community. Recently, researchers started to analyze metagenomes using reference-free methods based on the analysis of oligonucleotides (k-mers) frequency spectrum previously applied to isolated genomes. However, little is known about their correlation with the existing approaches for metagenomic feature extraction, as well as the limits of applicability. Here we evaluated a metagenomic pairwise dissimilarity measure based on short k-mer spectrum using the example of human gut microbiota, a biomedically significant object of study.ResultsWe developed a method for calculating pairwise dissimilarity (beta-diversity) of “shotgun” metagenomes based on short k-mer spectra (5≤k≤11). The method was validated on simulated metagenomes and further applied to a large collection of human gut metagenomes from the populations of the world (n=281). The k-mer spectrum-based measure was found to behave similarly to one based on mapping to a reference gene catalog, but different from one using a genome catalog. This difference turned out to be associated with a significant presence of viral reads in a number of metagenomes. Simulations showed limited impact of bacterial genetic variability as well as sequencing errors on k-mer spectra. Specific differences between the datasets from individual populations were identified.ConclusionsOur approach allows rapid estimation of pairwise dissimilarity between metagenomes. Though we applied this technique to gut microbiota, it should be useful for arbitrary metagenomes, even metagenomes with novel microbiota. Dissimilarity measure based on k-mer spectrum provides a wider perspective in comparison with the ones based on the alignment against reference sequence sets. It helps not to miss possible outstanding features of metagenomic composition, particularly related to the presence of an unknown bacteria, virus or eukaryote, as well as to technical artifacts (sample contamination, reads of non-biological origin, etc.) at the early stages of bioinformatic analysis. Our method is complementary to reference-based approaches and can be easily integrated into metagenomic analysis pipelines.Electronic supplementary materialThe online version of this article (doi:10.1186/s12859-015-0875-7) contains supplementary material, which is available to authorized users.

Highlights

A rapidly increasing flow of genomic data requires the development of efficient methods for obtaining its compact representation
To compare the k-mer based metagenomic beta-diversity measure with traditional reference-based methods we conducted a series of computational experiments on simulated and real data
The method was applied for the analysis of a group of real human gut metagenomes sequenced in two large-scale projects: China population (n = 152) [28] and HMP (n = 129) [27]

Summary

Introduction

A rapidly increasing flow of genomic data requires the development of efficient methods for obtaining its compact representation. We evaluated a metagenomic pairwise dissimilarity measure based on short k-mer spectrum using the example of human gut microbiota, a biomedically significant object of study. Advent of the next-generation sequencing allowed performing genomic analysis of samples obtained directly from the environment. Such an approach provides data for an extensive quantitative examination of the microbial community structure, including uncultivable and previously undiscovered components. One of the common steps in metagenomic study is calculation of pairwise dissimilarity between the samples (beta-diversity) [6]. In large-scale studies involving tens and hundreds of metagenomes, critical requirements in beta-diversity analysis include high algorithm performance and low memory usage

Methods

Results

Discussion

Conclusion

Talk to us

Join us for a 30 min session where you can share your feedback and ask us any queries you have

Schedule a call

Journal: BMC Bioinformatics	Publication Date: Jan 16, 2016
Citations: 110	License type: CC BY 4.0

R Discovery Prime

R Discovery Prime

Assessment of k-mer spectrum applicability for metagenomic dissimilarity analysis.

Abstract

Highlights

Summary

Talk to us

Similar Papers

More From: BMC Bioinformatics

Lead the way for us

Similar Papers

Genomic DNA k-mer spectra: models and modalities
Benny Chor ... Tim Massingham
Genome Biology | VOL. 10
Benny Chor, et. al.Benny Chor ... Tim Massingham
01 Jan 2009
Genome Biology | VOL. 10

Genomic DNA k-mer Spectra: Models and Modalities
Benny Chor ... Yaron Levy
-
Benny Chor, et. al.Benny Chor ... Yaron Levy
01 Jan 2009
01 Jan 2009

You Lose Some, You Win Some: Weight Loss Induces Microbiota and Metabolite Shifts
Elisabeth M Bik
EBioMedicine | VOL. 2
Elisabeth M BikElisabeth M Bik
30 Jul 2015
EBioMedicine | VOL. 2

Comparative Metagenomic Analysis of Chicken Gut Microbial Community, Function, and Resistome to Evaluate Noninvasive and Cecal Sampling Resources.
Kelang Kang ... Shourong Shi
Animals | VOL. 11
Kelang Kang, et. al.Kelang Kang ... Shourong Shi
09 Jun 2021
Animals | VOL. 11

Editage

Paperpal

R Discovery

Mind the Graph

R Discovery Prime

R Discovery Prime

Assessment of k-mer spectrum applicability for metagenomic dissimilarity analysis.

Abstract

Highlights

Summary

Talk to us

Similar Papers

More From: BMC Bioinformatics