Accelerate Literature Icon
Want to do a literature review? Try our new Literature Review workflow

Reducing the Effects of PCR Amplification and Sequencing Artifacts on 16S rRNA-Based Studies

  • Abstract
  • Highlights & Summary
  • PDF
  • Literature Map
  • Similar Papers
Abstract
Translate article icon Translate Article Star icon

The advent of next generation sequencing has coincided with a growth in interest in using these approaches to better understand the role of the structure and function of the microbial communities in human, animal, and environmental health. Yet, use of next generation sequencing to perform 16S rRNA gene sequence surveys has resulted in considerable controversy surrounding the effects of sequencing errors on downstream analyses. We analyzed 2.7×106 reads distributed among 90 identical mock community samples, which were collections of genomic DNA from 21 different species with known 16S rRNA gene sequences; we observed an average error rate of 0.0060. To improve this error rate, we evaluated numerous methods of identifying bad sequence reads, identifying regions within reads of poor quality, and correcting base calls and were able to reduce the overall error rate to 0.0002. Implementation of the PyroNoise algorithm provided the best combination of error rate, sequence length, and number of sequences. Perhaps more problematic than sequencing errors was the presence of chimeras generated during PCR. Because we knew the true sequences within the mock community and the chimeras they could form, we identified 8% of the raw sequence reads as chimeric. After quality filtering the raw sequences and using the Uchime chimera detection program, the overall chimera rate decreased to 1%. The chimeras that could not be detected were largely responsible for the identification of spurious operational taxonomic units (OTUs) and genus-level phylotypes. The number of spurious OTUs and phylotypes increased with sequencing effort indicating that comparison of communities should be made using an equal number of sequences. Finally, we applied our improved quality-filtering pipeline to several benchmarking studies and observed that even with our stringent data curation pipeline, biases in the data generation pipeline and batch effects were observed that could potentially confound the interpretation of microbial community data.

Similar Papers
  • PDF Download Icon
  • Preprint Article
  • Cite Count Icon 9
  • 10.7287/peerj.preprints.2196v3
Optimization of 16S amplicon analysis using mock communities: implications for estimating community diversity
  • Oct 26, 2016
  • Andrew Krohn + 5 more

The diversity of complex microbial communities can be rapidly assessed by high-throughput DNA sequencing of marker gene (e.g., 16S) PCR amplicon pools, often yielding many thousands of DNA sequences per sample. However, analysis of such community amplicon sequencing data requires multiple computational steps which affect the outcome of a final data set. Here we use mock communities to describe the effects of parameter adjustments for raw sequence quality filtering, picking operational taxonomic units (OTUs), taxonomic assignment, and OTU table filtering as implemented in the popular microbial ecology analysis package, QIIME 1.9.1. We demonstrate a workflow optimization based upon this exploration, which we also apply to environmental samples. We found that quality filtering of raw data and filtering of OTU tables had large effects on observed OTU diversity. While all taxonomy assignment programs performed with similar accuracy, an appropriate choice of similarity threshold for defining OTUs depended on the method used for OTU picking. Our “default” analysis in QIIME overestimated mock community OTU diversity by at least a factor of ten. Our optimized analysis correctly characterized mock community taxonomic composition and improved the OTU diversity estimate, reducing overestimation to a factor of about two. Though observed relative abundances of mock community member taxa were approximately correct, most were still represented by multiple OTUs. Low-frequency OTUs conspecific to constituent mock community taxa were characterized by multiple substitution and indel errors and the presence of a low-quality base call resulting in sequence truncation during quality filtering. Low-quality base calls were observed at “G” positions most of the time, and were also associated with a preceding “TTT” trinucleotide motif. Environmental diversity estimates were reduced by about 40% from 2508 to 1533 OTUs when comparing output from the default and optimized workflows. We attribute this reduction in observed diversity to the removal of erroneous sequences from the data set. Our results indicate that both strict quality filtering of raw sequencing data and careful filtering of raw OTU tables are important steps for accurately estimating microbial community diversity.

  • Preprint Article
  • Cite Count Icon 2
  • 10.7287/peerj.preprints.778v2
Sequencing 16S rRNA gene fragments using the PacBio SMRT DNA sequencing system
  • Feb 19, 2016
  • Patrick D Schloss + 4 more

Over the past 10 years, microbial ecologists have largely abandoned sequencing 16S rRNA genes by the Sanger sequencing method and have instead adopted highly parallelized sequencing platforms. These new platforms, such as 454 and Illumina's MiSeq, have allowed researchers to obtain millions of high quality, but short sequences. The result of the added sequencing depth has been significant improvements in experimental design. The tradeoff has been the decline in the number of full-length reference sequences that are deposited into databases. To overcome this problem, we tested the ability of the PacBio Single Molecule, Real-Time (SMRT) DNA sequencing platform to generate sequence reads from the 16S rRNA gene. We generated sequencing data from the V4, V3-V5, V1-V3, V1-V5, V1-V6, and V1-V9 variable regions from within the 16S rRNA gene using DNA from a synthetic mock community and natural samples collected from human feces, mouse feces, and soil. The mock community allowed us to assess the actual sequencing error rate and how that error rate changed when different curation methods were applied. We developed a simple method based on sequence characteristics and quality scores to reduce the observed error rate for the V1-V9 region from 0.69 to 0.027%. This error rate is comparable to what has been observed for the shorter reads generated by 454 and Illumina's MiSeq sequencing platforms. Although the per base sequencing cost is still significantly more than that of MiSeq, the prospect of supplementing reference databases with full-length sequences from organisms below the limit of detection from the Sanger approach is exciting.

  • PDF Download Icon
  • Research Article
  • Cite Count Icon 233
  • 10.7717/peerj.1869
Sequencing 16S rRNA gene fragments using the PacBio SMRT DNA sequencing system.
  • Mar 28, 2016
  • PeerJ
  • Patrick D Schloss + 4 more

Over the past 10 years, microbial ecologists have largely abandoned sequencing 16S rRNA genes by the Sanger sequencing method and have instead adopted highly parallelized sequencing platforms. These new platforms, such as 454 and Illumina’s MiSeq, have allowed researchers to obtain millions of high quality but short sequences. The result of the added sequencing depth has been significant improvements in experimental design. The tradeoff has been the decline in the number of full-length reference sequences that are deposited into databases. To overcome this problem, we tested the ability of the PacBio Single Molecule, Real-Time (SMRT) DNA sequencing platform to generate sequence reads from the 16S rRNA gene. We generated sequencing data from the V4, V3–V5, V1–V3, V1–V5, V1–V6, and V1–V9 variable regions from within the 16S rRNA gene using DNA from a synthetic mock community and natural samples collected from human feces, mouse feces, and soil. The mock community allowed us to assess the actual sequencing error rate and how that error rate changed when different curation methods were applied. We developed a simple method based on sequence characteristics and quality scores to reduce the observed error rate for the V1–V9 region from 0.69 to 0.027%. This error rate is comparable to what has been observed for the shorter reads generated by 454 and Illumina’s MiSeq sequencing platforms. Although the per base sequencing cost is still significantly more than that of MiSeq, the prospect of supplementing reference databases with full-length sequences from organisms below the limit of detection from the Sanger approach is exciting.

  • PDF Download Icon
  • Research Article
  • Cite Count Icon 466
  • 10.1371/journal.pone.0043093
PCR biases distort bacterial and archaeal community structure in pyrosequencing datasets.
  • Aug 15, 2012
  • PLoS ONE
  • Ameet J Pinto + 1 more

As 16S rRNA gene targeted massively parallel sequencing has become a common tool for microbial diversity investigations, numerous advances have been made to minimize the influence of sequencing and chimeric PCR artifacts through rigorous quality control measures. However, there has been little effort towards understanding the effect of multi-template PCR biases on microbial community structure. In this study, we used three bacterial and three archaeal mock communities consisting of, respectively, 33 bacterial and 24 archaeal 16S rRNA gene sequences combined in different proportions to compare the influences of (1) sequencing depth, (2) sequencing artifacts (sequencing errors and chimeric PCR artifacts), and (3) biases in multi-template PCR, towards the interpretation of community structure in pyrosequencing datasets. We also assessed the influence of each of these three variables on α- and β-diversity metrics that rely on the number of OTUs alone (richness) and those that include both membership and the relative abundance of detected OTUs (diversity). As part of this study, we redesigned bacterial and archaeal primer sets that target the V3–V5 region of the 16S rRNA gene, along with multiplexing barcodes, to permit simultaneous sequencing of PCR products from the two domains. We conclude that the benefits of deeper sequencing efforts extend beyond greater OTU detection and result in higher precision in β-diversity analyses by reducing the variability between replicate libraries, despite the presence of more sequencing artifacts. Additionally, spurious OTUs resulting from sequencing errors have a significant impact on richness or shared-richness based α- and β-diversity metrics, whereas metrics that utilize community structure (including both richness and relative abundance of OTUs) are minimally affected by spurious OTUs. However, the greatest obstacle towards accurately evaluating community structure are the errors in estimated mean relative abundance of each detected OTU due to biases associated with multi-template PCR reactions.

  • Research Article
  • Cite Count Icon 308
  • 10.1128/aem.02206-14
Performance comparison of Illumina and ion torrent next-generation sequencing platforms for 16S rRNA-based bacterial community profiling.
  • Sep 26, 2014
  • Applied and Environmental Microbiology
  • Stephen J Salipante + 8 more

High-throughput sequencing of the taxonomically informative 16S rRNA gene provides a powerful approach for exploring microbial diversity. Here we compare the performances of two common "benchtop" sequencing platforms, Illumina MiSeq and Ion Torrent Personal Genome Machine (PGM), for bacterial community profiling by 16S rRNA (V1-V2) amplicon sequencing. We benchmarked performance by using a 20-organism mock bacterial community and a collection of primary human specimens. We observed comparatively higher error rates with the Ion Torrent platform and report a pattern of premature sequence truncation specific to semiconductor sequencing. Read truncation was dependent on both the directionality of sequencing and the target species, resulting in organism-specific biases in community profiles. We found that these sequencing artifacts could be minimized by using bidirectional amplicon sequencing and an optimized flow order on the Ion Torrent platform. Results of bacterial community profiling performed on the mock community and a collection of 18 human-derived microbiological specimens were generally in good agreement for both platforms; however, in some cases, results differed significantly. Disparities could be attributed to the failure to generate full-length reads for particular organisms on the Ion Torrent platform, organism-dependent differences in sequence error rates affecting classification of certain species, or some combination of these factors. This study demonstrates the potential for differential bias in bacterial community profiles resulting from the choice of sequencing platform alone.

  • Research Article
  • Cite Count Icon 170
  • 10.1128/aem.00342-13
Distribution-Based Clustering: Using Ecology To Refine the Operational Taxonomic Unit
  • Aug 23, 2013
  • Applied and Environmental Microbiology
  • Sarah P Preheim + 4 more

16S rRNA sequencing, commonly used to survey microbial communities, begins by grouping individual reads into operational taxonomic units (OTUs). There are two major challenges in calling OTUs: identifying bacterial population boundaries and differentiating true diversity from sequencing errors. Current approaches to identifying taxonomic groups or eliminating sequencing errors rely on sequence data alone, but both of these activities could be informed by the distribution of sequences across samples. Here, we show that using the distribution of sequences across samples can help identify population boundaries even in noisy sequence data. The logic underlying our approach is that bacteria in different populations will often be highly correlated in their abundance across different samples. Conversely, 16S rRNA sequences derived from the same population, whether slightly different copies in the same organism, variation of the 16S rRNA gene within a population, or sequences generated randomly in error, will have the same underlying distribution across sampled environments. We present a simple OTU-calling algorithm (distribution-based clustering) that uses both genetic distance and the distribution of sequences across samples and demonstrate that it is more accurate than other methods at grouping reads into OTUs in a mock community. Distribution-based clustering also performs well on environmental samples: it is sensitive enough to differentiate between OTUs that differ by a single base pair yet predicts fewer overall OTUs than most other methods. The program can decrease the total number of OTUs with redundant information and improve the power of many downstream analyses to describe biologically relevant trends.

  • Research Article
  • 10.1158/1538-7445.am2019-441
Abstract 441: The stochastic nature of errors in next-generation sequencing of circulating cell-free DNA
  • Jul 1, 2019
  • Cancer Research
  • Hunter R Underhill + 6 more

Purpose: Challenges with distinguishing circulating tumor DNA from next-generation sequencing (NGS) artifacts limits variant searches to established solid tumor mutations. Identifying the source(s) of errors associated with the NGS analytics of circulating cell-free DNA (ccfDNA) would enable the determination of an optimal strategy for eliminating noise and broaden ccfDNA clinical applications. Methods: Buffy coat DNA and ccfDNA were isolated from seven healthy adults. For each participant, a single buffy coat DNA library was generated using duplex adapters (dual unique molecular identifiers [UMIs], dual index), while two ccfDNA libraries were separately produced - one library with singleton adapters (single UMI, single index) and one library with duplex adapters. The assignment of a UMI to each template DNA molecule prior to library formation reduces false positives. A family is a set of DNA amplicons (PCR duplicates) with the same UMI. Representing a family with a single consensus sequence reduces PCR errors and sequencing artifacts. Duplex adapters have been developed to abrogate the early PCR errors that beset singleton adapters. Results: The error rate using duplex adapters was significantly lower by 26.4±5.9% (P < 0.001) compared to singleton adapters at family size ≥2, where family size is defined as the number of PCR duplicates that yield a single consensus sequence. Due to the persistence of noise in both singleton and duplex adapters even at large family sizes (i.e., ≥10), we explored potential sources of the residual error. Noise in ccfDNA due to effects from clonal hematopoiesis of indeterminate potential (CHIP) accounted for <4% of error. Removing locations with errors present in all seven samples (i.e., highly patterned error likely due to regions difficult to sequence, align, or both) reduced noise in the duplex and singleton adapters by 18.7±5.1% (P < 0.001) and 14.5±3.1% (P < 0.001), respectively. Finally, we explored the effects of stochastic noise as a source of error. A complete replicate with duplex adapters was generated beginning from the source ccfDNA and library preparation. Using replicate data, the error rate was reduced by an additional 59.9±4.3% (P < 0.001) for the duplex adapters at family size ≥2. Using duplex adapters, accounting for CHIP artifacts, removing locations with highly patterned errors, and including replicate data reduced error by 85.4±4.6% (P < 0.001) compared to the error rate for singleton adapters at family size ≥2. Error continued to decline with each family size increment. Conclusion: Early stochastic PCR errors are a principal source of NGS noise that persist despite duplex molecular barcoding and after removal of patterned errors. Replicates are necessary to eliminate noise and their use in NGS analytics may broaden ccfDNA applications particularly in pre-metastatic and recurrent solid-tumor malignancies by enabling untargeted variant investigations. Citation Format: Hunter R. Underhill, Preetida J. Bhetariya, Sabine Hellwig, David A. Nix, Carrie L. Fuertes, Gabor T. Marth, Mary P. Bronner. The stochastic nature of errors in next-generation sequencing of circulating cell-free DNA [abstract]. In: Proceedings of the American Association for Cancer Research Annual Meeting 2019; 2019 Mar 29-Apr 3; Atlanta, GA. Philadelphia (PA): AACR; Cancer Res 2019;79(13 Suppl):Abstract nr 441.

  • PDF Download Icon
  • Research Article
  • Cite Count Icon 431
  • 10.1128/mspheredirect.00073-17
OptiClust, an Improved Method for Assigning Amplicon-Based Sequence Data to Operational Taxonomic Units
  • Mar 8, 2017
  • mSphere
  • Sarah L Westcott + 1 more

Assignment of 16S rRNA gene sequences to operational taxonomic units (OTUs) is a computational bottleneck in the process of analyzing microbial communities. Although this has been an active area of research, it has been difficult to overcome the time and memory demands while improving the quality of the OTU assignments. Here, we developed a new OTU assignment algorithm that iteratively reassigns sequences to new OTUs to optimize the Matthews correlation coefficient (MCC), a measure of the quality of OTU assignments. To assess the new algorithm, OptiClust, we compared it to 10 other algorithms using 16S rRNA gene sequences from two simulated and four natural communities. Using the OptiClust algorithm, the MCC values averaged 15.2 and 16.5% higher than the OTUs generated when we used the average neighbor and distance-based greedy clustering with VSEARCH, respectively. Furthermore, on average, OptiClust was 94.6 times faster than the average neighbor algorithm and just as fast as distance-based greedy clustering with VSEARCH. An empirical analysis of the efficiency of the algorithms showed that the time and memory required to perform the algorithm scaled quadratically with the number of unique sequences in the data set. The significant improvement in the quality of the OTU assignments over previously existing methods will significantly enhance downstream analysis by limiting the splitting of similar sequences into separate OTUs and merging of dissimilar sequences into the same OTU. The development of the OptiClust algorithm represents a significant advance that is likely to have numerous other applications. IMPORTANCE The analysis of microbial communities from diverse environments using 16S rRNA gene sequencing has expanded our knowledge of the biogeography of microorganisms. An important step in this analysis is the assignment of sequences into taxonomic groups based on their similarity to sequences in a database or based on their similarity to each other, irrespective of a database. In this study, we present a new algorithm for the latter approach. The algorithm, OptiClust, seeks to optimize a metric of assignment quality by shuffling sequences between taxonomic groups. We found that OptiClust produces more robust assignments and does so in a rapid and memory-efficient manner. This advance will allow for a more robust analysis of microbial communities and the factors that shape them.

  • Research Article
  • Cite Count Icon 16
  • 10.1111/nph.13851
Data processing can mask biology: towards better reporting of fungal barcoding data?
  • Jan 28, 2016
  • New Phytologist
  • Marc‐André Selosse + 2 more

Data processing can mask biology: towards better reporting of fungal barcoding data?

  • Abstract
  • Cite Count Icon 4
  • 10.1186/gb-2010-11-s1-i19
The rare biosphere: sorting out fact from fiction
  • Jan 1, 2010
  • Genome Biology
  • Mitchell L Sogin + 4 more

Over the past 25 years, microbiologists have employed the occurrence of DNA sequences as proxies for the presence of different kinds of organisms in microbial communities. These culture-independent investigations described new dimensions of diversity, identified novel candidate phyla, and redefined habitable ranges for single-cell organisms. The recent introduction of massively-parallel sequencing technology significantly increased estimates of microbial diversity from molecular-based studies. Matching SSU pyrotags to a reference rRNA database or clustering tags in a taxon-independent manner to identify Operational Taxonomic Units (OTUs) suggests that taxonomic richness in marine, terrestrial, and both the human and mouse microbiomes exceeds all prior estimates of microbial diversity. The occurrence of rare sequences in these data sets correspond to low abundance taxa that comprise the 'rare biosphere'. Numerous theories and mechanisms that could account for the existence and persistence of rare biosphere members compete with explanations that invoke sequencing or clustering artifacts. Even with sequencing error rates below 0.005 per nucleotide position, the common method of generating OTUs (i.e. multiple sequence alignment and complete-linkage clustering) significantly increases the number of predicted OTUs and inflates richness estimates. The use of a novel Single Linkage Preclustering (SLP) strategy applied to short hypervariable regions of ribosomal RNAs accurately identified the predicted complexity of 'mock' microbial communities with a known number of rRNA operons. The strategy initially identifies sequences that are likely to have arisen by error using nearest neighbor clustering of pairwise sequence distances. The most abundant sequence for each precluster and the number of sequences in the precluster define inputs to average neighbor clustering using MOTHUR. When applied to sequences obtained from multiple microbial communities, the OTU-based descriptions of microbial population structures under different ecological regimes, and the global distribution patterns of OTUs reinforce credibility of the 'rare biosphere' as revealed through deep sequencing efforts.

  • Research Article
  • 10.1158/1557-3265.liqbiop20-a57
Abstract A57: Uncovering instrument errors in next-generation sequencing by CleanDeepSeq2
  • Jun 1, 2020
  • Clinical Cancer Research
  • Eric Davis + 8 more

Liquid biopsy holds great promise in noninvasive diagnosis of cancers through detecting minute amounts of cell-free DNA released from cancer cells in non-solid biologic tissue such as peripheral blood. A critical bottleneck in developing liquid biopsy methods is the limited accuracy of current next-generation sequencing technology (NGS), evidenced by its high error rate (0.1%-1%, as of 2018). Through mathematical modeling of NGS errors, we have recently published a method to computationally suppress the current NGS error rate to between 10−5 and 10−4, two orders of magnitude lower than general reports. However, this error rate is a product of both PCR errors and instrument (i.e., sequencer) errors, and it is currently unknown how to separate these error sources. In this work, we developed a novel computational algorithm to precisely measure the errors caused by sequencers. By using 12 publicly available datasets from 10 sequencing centers (in America, Europe, and Asia), we discovered highly reproducible patterns of sequencer errors, including: 1) the overall sequencer error rate is 10−5; 2) at the flow-cell level, error rates are elevated in the bottom surface; 3) almost all flow cells have a small fraction of random tiles with a dramatically elevated error rate; 4) the elevated error rates appear to be enriched in some reaction cycles; 5) removal of these reaction cycles yields 5-fold lower error rates at some genomic loci, so that A>C, A>T, and C>G error types have error rates close to 10−6; and 6) sequencer errors have a pattern markedly distinct from PCR errors. We have implemented the above observations into a general-purpose algorithm, termed CleanDeepSeq2, to computationally suppress sequencer errors and to also effectively monitor sequencer anomalies. CleanDeepSeq2 was engineered for efficiency so that a dataset with ultra-deep sequencing (1,000,000X depth) can be processed in 1.5N minutes on a single CPU core, where N is the number of target regions. Similarly, WES (100X) and WGS (~30X) datasets can be processed in under 1 CPU hour in order to monitor instrument performance. Overall, we have developed a computational method that for the first time enabled precise measurement of sequencer errors. Our study revealed novel insights on sequencer errors that can lead to improved instrumentation, NGS chemistry, and ultimately higher DNA sequencing fidelity. In addition, our developed software can efficiently suppress sequencer errors in addition to previously discovered error sources. Citation Format: Eric Davis, Rain Sun, Ying Shao, Yanling Liu, Heather L. Mulder, Stephen V. Rice, John Easton, Jinghui Zhang, Xiaotu Ma. Uncovering instrument errors in next-generation sequencing by CleanDeepSeq2 [abstract]. In: Proceedings of the AACR Special Conference on Advances in Liquid Biopsies; Jan 13-16, 2020; Miami, FL. Philadelphia (PA): AACR; Clin Cancer Res 2020;26(11_Suppl):Abstract nr A57.

  • Conference Article
  • 10.1158/1538-7445.sabcs18-441
Abstract 441: The stochastic nature of errors in next-generation sequencing of circulating cell-free DNA
  • Jul 1, 2019
  • Clinical Research (Excluding Clinical Trials)
  • Hunter R Underhill + 6 more

Purpose: Challenges with distinguishing circulating tumor DNA from next-generation sequencing (NGS) artifacts limits variant searches to established solid tumor mutations. Identifying the source(s) of errors associated with the NGS analytics of circulating cell-free DNA (ccfDNA) would enable the determination of an optimal strategy for eliminating noise and broaden ccfDNA clinical applications. Methods: Buffy coat DNA and ccfDNA were isolated from seven healthy adults. For each participant, a single buffy coat DNA library was generated using duplex adapters (dual unique molecular identifiers [UMIs], dual index), while two ccfDNA libraries were separately produced - one library with singleton adapters (single UMI, single index) and one library with duplex adapters. The assignment of a UMI to each template DNA molecule prior to library formation reduces false positives. A family is a set of DNA amplicons (PCR duplicates) with the same UMI. Representing a family with a single consensus sequence reduces PCR errors and sequencing artifacts. Duplex adapters have been developed to abrogate the early PCR errors that beset singleton adapters. Results: The error rate using duplex adapters was significantly lower by 26.4±5.9% (P Conclusion: Early stochastic PCR errors are a principal source of NGS noise that persist despite duplex molecular barcoding and after removal of patterned errors. Replicates are necessary to eliminate noise and their use in NGS analytics may broaden ccfDNA applications particularly in pre-metastatic and recurrent solid-tumor malignancies by enabling untargeted variant investigations. Citation Format: Hunter R. Underhill, Preetida J. Bhetariya, Sabine Hellwig, David A. Nix, Carrie L. Fuertes, Gabor T. Marth, Mary P. Bronner. The stochastic nature of errors in next-generation sequencing of circulating cell-free DNA [abstract]. In: Proceedings of the American Association for Cancer Research Annual Meeting 2019; 2019 Mar 29-Apr 3; Atlanta, GA. Philadelphia (PA): AACR; Cancer Res 2019;79(13 Suppl):Abstract nr 441.

  • Conference Article
  • Cite Count Icon 1
  • 10.1145/3233547.3233612
HmmUFOtu
  • Aug 15, 2018
  • Qi Zheng + 3 more

Over the last decade, joint advances in next-generation sequencing technology and bioinformatics pipelines have dramatically improved our understanding of host-associated and environmental microbiota. Standard microbiome community analysis typically involves amplicon sequencing of the prokaryotic 16S rRNA gene. These sequences are then clustered into operational taxonomic units (OTUs) for downstream diversity analyses, but also to reduce computational burden and allow for rapid analysis of datasets. Taxonomy is then assigned to all reads of an OTU, based on the assignment of a representative read. Although straightforward in principle, present methods often rely on heuristics while constructing (or "picking") OTUs to avoid computationally expensive algorithms, and ignore the prior knowledge of microbial phylogeny to further reduce the computational complexity. Here, we present HmmUFOtu, a novel tool for processing 16S rRNA sequences that addresses major limitations of current OTU picking and taxonomy assignment methods. HmmUFOtu relies on rapid per-read phylogenetic placement, followed by OTU picking and taxonomic assignment based on the phylogeny of known taxa. By benchmarking on simulated, mock community, and real datasets, we show that HmmUFOtu achieves high assignment accuracy, sensitivity, specificity and precision, even at species-level resolution. Compared to standard pipelines, HmmUFOtu more accurately recapitulates community diversity and composition. HmmUFOtu can perform taxonomic assignment in a species-resolution reference tree with ~ 200,000 nodes for 1 million 16S sequencing reads within 6 hours on a modest Linux workstation with 16 processors and 32 GB RAM. HmmUFOtu is written in C++98 and freely available at https://github.com/Grice-Lab/HmmUFOtu/.

  • PDF Download Icon
  • Research Article
  • Cite Count Icon 453
  • 10.1128/msphere.01202-20
Primer, Pipelines, Parameters: Issues in 16S rRNA Gene Sequencing.
  • Feb 24, 2021
  • mSphere
  • Isabel Abellan-Schneyder + 7 more

ABSTRACTShort-amplicon 16S rRNA gene sequencing is currently the method of choice for studies investigating microbiomes. However, comparative studies on differences in procedures are scarce. We sequenced human stool samples and mock communities with increasing complexity using a variety of commonly used protocols. Short amplicons targeting different variable regions (V-regions) or ranges thereof (V1-V2, V1-V3, V3-V4, V4, V4-V5, V6-V8, and V7-V9) were investigated for differences in the composition outcome due to primer choices. Next, the influence of clustering (operational taxonomic units [OTUs], zero-radius OTUs [zOTUs], and amplicon sequence variants [ASVs]), different databases (GreenGenes, the Ribosomal Database Project, Silva, the genomic-based 16S rRNA Database, and The All-Species Living Tree), and bioinformatic settings on taxonomic assignment were also investigated. We present a systematic comparison across all typically used V-regions using well-established primers. While it is known that the primer choice has a significant influence on the resulting microbial composition, we show that microbial profiles generated using different primer pairs need independent validation of performance. Further, comparing data sets across V-regions using different databases might be misleading due to differences in nomenclature (e.g., Enterorhabdus versus Adlercreutzia) and varying precisions in classification down to genus level. Overall, specific but important taxa are not picked up by certain primer pairs (e.g., Bacteroidetes is missed using primers 515F-944R) or due to the database used (e.g., Acetatifactor in GreenGenes and the genomic-based 16S rRNA Database). We found that appropriate truncation of amplicons is essential and different truncated-length combinations should be tested for each study. Finally, specific mock communities of sufficient and adequate complexity are highly recommended.IMPORTANCE In 16S rRNA gene sequencing, certain bacterial genera were found to be underrepresented or even missing in taxonomic profiles when using unsuitable primer combinations, outdated reference databases, or inadequate pipeline settings. Concerning the last, quality thresholds as well as bioinformatic settings (i.e., clustering approach, analysis pipeline, and specific adjustments such as truncation) are responsible for a number of observed differences between studies. Conclusions drawn by comparing one data set to another (e.g., between publications) appear to be problematic and require independent cross-validation using matching V-regions and uniform data processing. Therefore, we highlight the importance of a thought-out study design including sufficiently complex mock standards and appropriate V-region choice for the sample of interest. The use of processing pipelines and parameters must be tested beforehand.

  • PDF Download Icon
  • Research Article
  • Cite Count Icon 37
  • 10.1371/journal.pone.0017034
Error and Error Mitigation in Low-Coverage Genome Assemblies
  • Feb 14, 2011
  • PLoS ONE
  • Melissa J Hubisz + 3 more

The recent release of twenty-two new genome sequences has dramatically increased the data available for mammalian comparative genomics, but twenty of these new sequences are currently limited to ∼2× coverage. Here we examine the extent of sequencing error in these 2× assemblies, and its potential impact in downstream analyses. By comparing 2× assemblies with high-quality sequences from the ENCODE regions, we estimate the rate of sequencing error to be 1–4 errors per kilobase. While this error rate is fairly modest, sequencing error can still have surprising effects. For example, an apparent lineage-specific insertion in a coding region is more likely to reflect sequencing error than a true biological event, and the length distribution of coding indels is strongly distorted by error. We find that most errors are contributed by a small fraction of bases with low quality scores, in particular, by the ends of reads in regions of single-read coverage in the assembly. We explore several approaches for automatic sequencing error mitigation (SEM), making use of the localized nature of sequencing error, the fact that it is well predicted by quality scores, and information about errors that comes from comparisons across species. Our automatic methods for error mitigation cannot replace the need for additional sequencing, but they do allow substantial fractions of errors to be masked or eliminated at the cost of modest amounts of over-correction, and they can reduce the impact of error in downstream phylogenomic analyses. Our error-mitigated alignments are available for download.

Save Icon
Up Arrow
Open/Close
Setting-up Chat
Loading Interface