Tag jumps illuminated--reducing sequence-to-sample misidentifications in metabarcoding studies.
Metabarcoding of environmental samples on second-generation sequencing platforms has rapidly become a valuable tool for ecological studies. A fundamental assumption of this approach is the reliance on being able to track tagged amplicons back to the samples from which they originated. In this study, we address the problem of sequences in metabarcoding sequencing outputs with false combinations of used tags (tag jumps). Unless these sequences can be identified and excluded from downstream analyses, tag jumps creating sequences with false, but already used tag combinations, can cause incorrect assignment of sequences to samples and artificially inflate diversity. In this study, we document and investigate tag jumping in metabarcoding studies on Illumina sequencing platforms by amplifying mixed-template extracts obtained from bat droppings and leech gut contents with tagged generic arthropod and mammal primers, respectively. We found that an average of 2.6% and 2.1% of sequences had tag combinations, which could be explained by tag jumping in the leech and bat diet study, respectively. We suggest that tag jumping can happen during blunt-ending of pools of tagged amplicons during library build and as a consequence of chimera formation during bulk amplification of tagged amplicons during library index PCR. We argue that tag jumping and contamination between libraries represents a considerable challenge for Illumina-based metabarcoding studies, and suggest measures to avoid false assignment of tag jumping-derived sequences to samples.
- Research Article
1
- 10.3897/aca.4.e65450
- Mar 4, 2021
- ARPHA Conference Abstracts
Labelling strategies in metabarcoding studies & how to ensure that nucleotide tags stay in place Metabarcoding of environmental DNA (eDNA) and DNA extracted from bulk specimen samples is a powerful tool in studies of ecological interactions, diet and biodiversity, as its labelling of amplicons allows high-throughput sequencing of taxonomically informative DNA sequences from many samples in parallel. The backbone of metabarcoding is the addition of sample-specific nucleotide identifiers to amplicons and then following sequencing using these to assign metabarcoding sequences to the samples they originated from. This allows the pooling of hundreds to thousands of samples before sequencing and thereby full utilisation of the capacity of high-throughput sequencing platforms. The nucleotide identifiers can be added both during the metabarcoding PCR and during library preparation, i.e. when amplicons are prepared for sequencing. There are three main strategies with which to achieve nucleotide labelling in metabarcoding studies. One commonly used strategy is the so-called tagged PCR approach in which DNA extracts are individually amplified with metabarcoding primers that carry sample-specific nucleotide tags at the 5’ end. The uniquely tagged products are then pooled and a library prepared on the pool of amplicons. However, tag‐jumps have been documented in this commonly used metabarcoding approach (Schnell et al. 2015). Tag-jumps cause nucleotide tags to switch between amplicons, resulting in occurrence of amplicons that carry different tags than originally applied. Sequences in the sequencing output that carry tag combinations not used in the study design are easily identified and excluded. However, sequences carrying incorrect, but already used, tag combinations will cause incorrect assignments of sequences to samples. This can - much to the detriment of metabarcoding studies - lead to false positives and artificial inflation of diversity in the samples (Schnell et al. 2015). The occurrence of tag-jumps has led to recommendations to only carry out metabarcoding PCR amplifications with primers carrying twin-tags to ensure that tag‐jumps cannot result in false assignments of sequences to samples (Schnell et al. 2015). However, this increases both cost and workload of metabarcoding studies. In a recently published article, we demonstrate a tag-jump free single-tube library preparation protocol for Illumina sequencing specifically designed for 5’ nucleotide tagged amplicons, the Tagsteady protocol (Carøe & Bohmann 2020). We designed the Tagsteady protocol to circumvent the two steps during library preparation of pools of 5ʹ nucleotide-tagged amplicons that had previously been suggested to cause tag-jumps; i) T4 DNA polymerase blunt-ending in the end-repair step, and ii) post-ligation PCR amplification of amplicon libraries. We used pools of twin‐tagged amplicons to investigate the effect of these two steps on the occurrence of tag‐jumps. Doing this, we demonstrated that blunt‐ending and post-ligation PCR, alone or together, can result in high proportions of tag-jumps, in our study up to ca. 49% of total sequences. The Tagsteady protocol where both these steps were left out resulted in tag‐jump levels comparable to background contamination (Carøe & Bohmann 2020). In our study, we encourage practitioners to avoid using T4 DNA polymerase blunt‐ending and post-ligation PCR in library preparation of 5’ nucleotide tagged amplicon pools, for example by using the Tagsteady protocol (Carøe & Bohmann 2020). This will enable efficient and cost-effective generation of metabarcoding data with correct assignment of sequences to samples. References Carøe C, Bohmann K (2020) Tagsteady: A metabarcoding library preparation protocol to avoid false assignment of sequences to samples. Molecular Ecology Resources, 20, 1620–1631. Schnell IB, Bohmann K, Gilbert MTP (2015) Tag jumps illuminated - reducing sequence-to-sample misidentifications in metabarcoding studies. Molecular Ecology Resources, 15, 1289–1303.
- Research Article
32
- 10.1111/1755-0998.12770
- Mar 8, 2018
- Molecular Ecology Resources
Different second-generation sequencing technologies may have taxon-specific biases when DNA metabarcoding prey in predator faeces. Our major objective was to examine differences in prey recovery from bat guano across two different sequencing workflows using the same faecal DNA extracts. We compared results between the Ion Torrent PGM and the Illumina MiSeq with similar library preparations and the same analysis pipeline. We focus on repeatability and provide an R Notebook in an effort towards transparency for future methodological improvements. Full documentation of each step enhances the accessibility of our analysis pipeline. We tagged DNA from insectivorous bat faecal samples, targeted the arthropod cytochrome c oxidase I minibarcode region and sequenced the product on both second-generation sequencing platforms. We developed an analysis pipeline with a high operational taxonomic unit (OTU) clustering threshold (i.e., ≥98.5%) followed by copy number filtering to avoid merging rare but genetically similar prey into the same OTUs. With this workflow, we detected 297 unique prey taxa, of which 74% were identified at the species level. Of these, 104 (35%) prey OTUs were detected by both platforms, 176 (59%) OTUs were detected by the Illumina MiSeq system only, and 17 (6%) OTUs were detected using the Ion Torrent system only. Costs were similar between platforms but the Illumina MiSeq recovered six times more reads and four additional insect orders than did Ion Torrent. The considerations we outline are particularly important for long-term ecological monitoring; a more standardized approach will facilitate comparisons between studies and allow faster recognition of changes within ecological communities.
- Research Article
26
- 10.3389/fpubh.2021.710985
- Aug 27, 2021
- Frontiers in Public Health
Fast and accurate identification of pathogens is an essential task in healthcare settings. Second-generation sequencing platforms such as Illumina have greatly expanded the capacity with which different organisms can be detected in hospital samples, and third-generation nanopore-driven sequencing devices such as Oxford Nanopore's minION have recently emerged as ideal sequencing platforms for routine healthcare surveillance due to their long-read capacity and high portability. Despite its great potential, protocols and analysis pipelines for nanopore sequencing are still being extensively validated. In this work, we assess the ability of nanopore sequencing to provide reliable community profiles based on 16S rRNA sequencing in comparison to traditional Illumina platforms using samples collected from Intensive Care Units of a hospital in Brazil. While our results demonstrate that lower throughputs may be a shortcoming of the method in more complex samples, we show that the use of single-use Flongle flowcells in nanopore sequencing runs can provide insightful information on the community composition in healthcare settings.
- Research Article
431
- 10.1128/mspheredirect.00073-17
- Mar 8, 2017
- mSphere
Assignment of 16S rRNA gene sequences to operational taxonomic units (OTUs) is a computational bottleneck in the process of analyzing microbial communities. Although this has been an active area of research, it has been difficult to overcome the time and memory demands while improving the quality of the OTU assignments. Here, we developed a new OTU assignment algorithm that iteratively reassigns sequences to new OTUs to optimize the Matthews correlation coefficient (MCC), a measure of the quality of OTU assignments. To assess the new algorithm, OptiClust, we compared it to 10 other algorithms using 16S rRNA gene sequences from two simulated and four natural communities. Using the OptiClust algorithm, the MCC values averaged 15.2 and 16.5% higher than the OTUs generated when we used the average neighbor and distance-based greedy clustering with VSEARCH, respectively. Furthermore, on average, OptiClust was 94.6 times faster than the average neighbor algorithm and just as fast as distance-based greedy clustering with VSEARCH. An empirical analysis of the efficiency of the algorithms showed that the time and memory required to perform the algorithm scaled quadratically with the number of unique sequences in the data set. The significant improvement in the quality of the OTU assignments over previously existing methods will significantly enhance downstream analysis by limiting the splitting of similar sequences into separate OTUs and merging of dissimilar sequences into the same OTU. The development of the OptiClust algorithm represents a significant advance that is likely to have numerous other applications. IMPORTANCE The analysis of microbial communities from diverse environments using 16S rRNA gene sequencing has expanded our knowledge of the biogeography of microorganisms. An important step in this analysis is the assignment of sequences into taxonomic groups based on their similarity to sequences in a database or based on their similarity to each other, irrespective of a database. In this study, we present a new algorithm for the latter approach. The algorithm, OptiClust, seeks to optimize a metric of assignment quality by shuffling sequences between taxonomic groups. We found that OptiClust produces more robust assignments and does so in a rapid and memory-efficient manner. This advance will allow for a more robust analysis of microbial communities and the factors that shape them.
- Research Article
42
- 10.1002/edn3.382
- Dec 18, 2022
- Environmental DNA
Metabarcoding of environmental DNA (eDNA) is a powerful tool for describing biodiversity, such as finding keystone species or detecting invasive species in environmental samples. Continuous improvements in the method and the advances in sequencing platforms over the last decade have meant this approach is now widely used in biodiversity sciences and biomonitoring. For its general use, the method hinges on a correct identification of taxa. However, past studies have shown how this crucially depends on important decisions during sampling, sample processing, and subsequent handling of sequencing data. With no clear consensus as to the best practice, particularly the latter has led to varied bioinformatic approaches and recommendations for data preparation and taxonomic identification. In this study, using a large freshwater fish eDNA sequence dataset, we compared the frequently used zero‐radius Operational Taxonomic Unit (zOTU) approach of our raw reads and assigned it taxonomically (i) in combination with publicly available reference sequences (open databases) or (ii) with an OSU (Operational Sequence Units) database approach, using a curated database of reference sequences generated from specimen barcoding (closed database). We show both approaches gave comparable results for common species. However, the commonalities between the approaches decreased with read abundance and were thus less reliable and not comparable for rare species. The success of the zOTU approach depended on the suitability, rather than the size, of a reference database. Contrastingly, the OSU approach used reliable DNA sequences and thus often enabled species‐level identifications, yet this resolution decreased with the recent phylogenetic age of the species. We show the need to include target group coverage, outgroups and full taxonomic annotation in reference databases to avoid misleading annotations that can occur when using short amplicon sizes as commonly used in eDNA metabarcoding studies. Finally, we make general suggestions to improve the construction and use of reference databases for metabarcoding studies in the future.
- Book Chapter
5
- 10.1016/bs.aecr.2018.06.002
- Jan 1, 2018
Bioinformatics for Biomonitoring: Species Detection and Diversity Estimates Across Next-Generation Sequencing Platforms
- Research Article
1395
- 10.1186/1471-2105-11-485
- Sep 27, 2010
- BMC Bioinformatics
BackgroundIllumina's second-generation sequencing platform is playing an increasingly prominent role in modern DNA and RNA sequencing efforts. However, rapid, simple, standardized and independent measures of run quality are currently lacking, as are tools to process sequences for use in downstream applications based on read-level quality data.ResultsWe present SolexaQA, a user-friendly software package designed to generate detailed statistics and at-a-glance graphics of sequence data quality both quickly and in an automated fashion. This package contains associated software to trim sequences dynamically using the quality scores of bases within individual reads.ConclusionThe SolexaQA package produces standardized outputs within minutes, thus facilitating ready comparison between flow cell lanes and machine runs, as well as providing immediate diagnostic information to guide the manipulation of sequence data for downstream analyses.
- Research Article
1
- 10.1002/ece3.71333
- May 1, 2025
- Ecology and evolution
Biodiversity monitoring using metabarcoding is now widely used as a routine environmental management tool. However, despite the rapid advancement of third-generation high-throughput sequencing platforms, there are limited studies assessing the most suitable tools and approaches for environmental metabarcoding studies. We tested the utility of Oxford Nanopore Technologies MinION sequencing for short-read amplicon sequencing of mitochondrial COI mini-barcodes from a known composition of arthropod species and compared its performance with more commonly used Illumina NovaSeq sequencing. The mock arthropod species assemblage allowed us to optimise a bioinformatic filtering pipeline to identify arthropod species using MinION long reads. Using this pipeline, we identified host species and diet composition by sequencing droppings collected from five individual Irish brown long-eared bats (Plecotus auritus) roosts. We showed that MinION data provided a similar taxonomic assignment to NovaSeq but only if the reference species barcode database was accurate and comprehensive. The P. auritus diet inferred was as expected based on previous morphological and Illumina metabarcoding studies. We showed that less sequencing depth, but a higher number of biological samples were necessary for complete species composition detection by MinION. A relatively simple bioinformatic filtering tool such as NanoPipe could adequately retrieve both host species and diet composition. The biggest standing challenge was the reference database format transferability and comprehensiveness. This pipeline can be used to guide future metabarcoding studies using nanopore sequencing to minimise the cost and effort while optimising results.
- Research Article
- 10.2139/ssrn.3899428
- Aug 4, 2021
- SSRN Electronic Journal
The META Tool Optimizes Metagenomic Analyses Across Sequencing Platforms and Classifiers
- Research Article
- 10.3389/fbinf.2022.969247
- Jan 6, 2023
- Frontiers in Bioinformatics
A major challenge in the field of metagenomics is the selection of the correct combination of sequencing platform and downstream metagenomic analysis algorithm, or “classifier”. Here, we present the Metagenomic Evaluation Tool Analyzer (META), which produces simulated data and facilitates platform and algorithm selection for any given metagenomic use case. META-generated in silico read data are modular, scalable, and reflect user-defined community profiles, while the downstream analysis is done using a variety of metagenomic classifiers. Reported results include information on resource utilization, time-to-answer, and performance. Real-world data can also be analyzed using selected classifiers and results benchmarked against simulations. To test the utility of the META software, simulated data was compared to real-world viral and bacterial metagenomic samples run on four different sequencers and analyzed using 12 metagenomic classifiers. Lastly, we introduce “META Score”: a unified, quantitative value which rates an analytic classifier’s ability to both identify and count taxa in a representative sample.
- Peer Review Report
- 10.7287/peerj.14616v0.1/reviews/1
- Jan 9, 2023
Background.In metabarcoding analyses, the taxonomic assignment is crucial to place sequencing data in biological and ecological contexts.This fundamental step depends on a reference database, which should have a good taxonomic coverage to avoid unassigned sequences.However, this goal is rarely achieved in many geographic regions and for several taxonomic groups.On the other hand, more is not necessarily better, as sequences in reference databases belonging to taxonomic groups out of the studied region/environment context might lead to false assignments. Methods.We investigated the effect of using several subsets of a cytochrome c oxidase subunit 1 (COI) reference database on taxonomic assignment.Published metabarcoding sequences from the Mediterranean Sea were assigned to taxa using COInr, which is a comprehensive, non-redundant and recent database of COI sequences obtained both from BOLD and NCBI, and two of its subsets: (i) all sequences except insects (COInr-WO-Insecta), which represent the overwhelming majority of COInr database, but are irrelevant for marine samples, and (ii) all sequences from taxonomic families present in the Mediterranean Sea (COInr-Med).Four different taxonomic assignation algorithms were employed in parallel to evaluate differences in their output and data consistency.Results.The reduction of the database to more specific custom subsets increased the number of unassigned sequences.Nevertheless, since most of them were incorrectly assigned by the less specific databases, this is a positive outcome.Moreover, the taxonomic resolution (the lowest taxonomic level to which a sequence is attributed) of several sequences tended to increase when using customized databases.These findings clearly indicated the need for customized databases adapted to each study.However, the very high proportion of unassigned sequences points to the need to enrich the local database with new barcodes specifically obtained from the studied region and/or taxonomic group.Including novel local barcodes to the COI database proved to be very profitable: by adding only 116 new barcodes sequenced in our laboratory, thus increasing the reference database by only 0.04%, we were able to improve the resolution for ca.0.6-1% of the Amplicon Sequence Variants (ASV).
- Peer Review Report
- 10.7287/peerj.14616v0.1/reviews/2
- Jan 9, 2023
Background.In metabarcoding analyses, the taxonomic assignment is crucial to place sequencing data in biological and ecological contexts.This fundamental step depends on a reference database, which should have a good taxonomic coverage to avoid unassigned sequences.However, this goal is rarely achieved in many geographic regions and for several taxonomic groups.On the other hand, more is not necessarily better, as sequences in reference databases belonging to taxonomic groups out of the studied region/environment context might lead to false assignments. Methods.We investigated the effect of using several subsets of a cytochrome c oxidase subunit 1 (COI) reference database on taxonomic assignment.Published metabarcoding sequences from the Mediterranean Sea were assigned to taxa using COInr, which is a comprehensive, non-redundant and recent database of COI sequences obtained both from BOLD and NCBI, and two of its subsets: (i) all sequences except insects (COInr-WO-Insecta), which represent the overwhelming majority of COInr database, but are irrelevant for marine samples, and (ii) all sequences from taxonomic families present in the Mediterranean Sea (COInr-Med).Four different taxonomic assignation algorithms were employed in parallel to evaluate differences in their output and data consistency.Results.The reduction of the database to more specific custom subsets increased the number of unassigned sequences.Nevertheless, since most of them were incorrectly assigned by the less specific databases, this is a positive outcome.Moreover, the taxonomic resolution (the lowest taxonomic level to which a sequence is attributed) of several sequences tended to increase when using customized databases.These findings clearly indicated the need for customized databases adapted to each study.However, the very high proportion of unassigned sequences points to the need to enrich the local database with new barcodes specifically obtained from the studied region and/or taxonomic group.Including novel local barcodes to the COI database proved to be very profitable: by adding only 116 new barcodes sequenced in our laboratory, thus increasing the reference database by only 0.04%, we were able to improve the resolution for ca.0.6-1% of the Amplicon Sequence Variants (ASV).
- Research Article
19
- 10.1002/edn3.255
- Nov 3, 2021
- Environmental DNA
How does the evolution of bioinformatics tools impact the biological interpretation of high‐throughput sequencing datasets? For eukaryotic metabarcoding studies, in particular, researchers often rely on tools originally developed for the analysis of 16S ribosomal RNA (rRNA) datasets. Such tools do not adequately account for the complexity of eukaryotic genomes, the ubiquity of intragenomic variation in eukaryotic metabarcoding loci, or the differential evolutionary rates observed across eukaryotic genes and taxa. Recently, metabarcoding workflows have shifted away from the use of operational taxonomic units (OTUs) toward delimitation of amplicon sequence variants (ASVs). We assessed how the choice of bioinformatics algorithm impacts the downstream biological conclusions that are drawn from eukaryotic 18S rRNA metabarcoding studies. We focused on four workflows including UCLUST and VSearch algorithms for OTU clustering, and DADA2 and Deblur algorithms for ASV delimitation. We used two 18S rRNA datasets to further evaluate whether dataset complexity had a major impact on the statistical trends and ecological metrics: a “high complexity” (HC) environmental dataset generated from community DNA in Arctic marine sediments, and a “low complexity” (LC) dataset representing individually barcoded nematodes. Our results indicate that ASV algorithms produce more biologically realistic metabarcoding outputs, with DADA2 being the most consistent and accurate pipeline regardless of dataset complexity. In contrast, OTU clustering algorithms inflate the metabarcoding‐derived estimates of biodiversity, consistently returning a high proportion of “rare” molecular operational taxonomic units (MOTUs) that appear to represent computational artifacts and sequencing errors. However, species‐specific MOTUs with high relative abundance are often recovered regardless of the bioinformatics approach. We also found high concordance across pipelines for downstream ecological analysis based on beta‐diversity and alpha‐diversity comparisons that utilize taxonomic assignment information. Analyses of LC datasets and rare MOTUs are especially sensitive to the choice of algorithms and better software tools may be needed to address these scenarios.
- Research Article
5
- 10.1002/edn3.70080
- Mar 1, 2025
- Environmental DNA (Hoboken, N.J.)
In metabarcoding studies, Linnaean taxonomy assignments of Operational Taxonomic Units (OTUs) or Amplicon Sequence Variants (ASVs) underpin many downstream bioinformatics analyses and ecological interpretations of environmental DNA (eDNA) datasets. However, public molecular databases (i.e., SILVA, EUKARYOME, BOLD) for most microbial metazoan phyla (nematodes, tardigrades, kinorhynchs, etc.) are sparsely populated, negatively impacting our ability to assign ecologically meaningful taxonomy to these understudied groups. Additionally, the choice of bioinformatics parameters and computational algorithms can further impact the accuracy of eDNA taxonomy assignments. Here, we use two in-silico datasets to show that taxonomy assignments using the 18S rRNA gene can be dramatically improved by curating Linnaean taxonomy strings associated with each reference sequence and closing phylogenetic gaps by improving taxon sampling. Using free-living nematodes as a case study, we applied two commonly used taxonomy assignment algorithms (BLAST+ and the QIIME2 Naïve Bayes classifier) across six iterations of the SILVA 138 reference database to evaluate the precision and accuracy of taxonomy assignments. The BLAST+ top hit with a 90% sequence similarity cutoff often returned the highest percentage of correctly assigned taxonomy at the genus level, and the QIIME2 Naïve Bayes classifier performed similarly well when paired with a reference database containing corrected taxonomy strings. Our results highlight the urgent need for phylogenetically-informed expansions of public reference databases (encompassing both genomes and common gene markers), focused on poorly sampled lineages which are now robustly recovered via eDNA metabarcoding approaches. Additional taxonomy curation efforts should be applied to popular reference databases such as SILVA, and taxon sampling could be rapidly improved by more frequent incorporation of newly published GenBank sequences linked to genus and/or species level identifications.
- Research Article
125
- 10.1101/gr.122747.111
- Jul 29, 2011
- Genome Research
Second-generation sequencing platforms have revolutionized the field of ancient DNA, opening access to complete genomes of past individuals and extinct species. However, these platforms are dependent on library construction and amplification steps that may result in sequences that do not reflect the original DNA template composition. This is particularly true for ancient DNA, where templates have undergone extensive damage post-mortem. Here, we report the results of the first "true single molecule sequencing" of ancient DNA. We generated 115.9 Mb and 76.9 Mb of DNA sequences from a permafrost-preserved Pleistocene horse bone using the Helicos HeliScope and Illumina GAIIx platforms, respectively. We find that the percentage of endogenous DNA sequences derived from the horse is higher among the Helicos data than Illumina data. This result indicates that the molecular biology tools used to generate sequencing libraries of ancient DNA molecules, as required for second-generation sequencing, introduce biases into the data that reduce the efficiency of the sequencing process and limit our ability to fully explore the molecular complexity of ancient DNA extracts. We demonstrate that simple modifications to the standard Helicos DNA template preparation protocol further increase the proportion of horse DNA for this sample by threefold. Comparison of Helicos-specific biases and sequence errors in modern DNA with those in ancient DNA also reveals extensive cytosine deamination damage at the 3' ends of ancient templates, indicating the presence of 3'-sequence overhangs. Our results suggest that paleogenomes could be sequenced in an unprecedented manner by combining current second- and third-generation sequencing approaches.