Accelerate Literature Icon
Want to do a literature review? Try our new Literature Review workflow

Trimmomatic: a flexible trimmer for Illumina sequence data

  • Abstract
  • Highlights & Summary
  • Literature Map
  • Similar Papers
Abstract
Translate article icon Translate Article Star icon

Motivation: Although many next-generation sequencing (NGS) read preprocessing tools already existed, we could not find any tool or combination of tools that met our requirements in terms of flexibility, correct handling of paired-end data and high performance. We have developed Trimmomatic as a more flexible and efficient preprocessing tool, which could correctly handle paired-end data.Results: The value of NGS read preprocessing is demonstrated for both reference-based and reference-free tasks. Trimmomatic is shown to produce output that is at least competitive with, and in many cases superior to, that produced by other tools, in all scenarios tested.Availability and implementation: Trimmomatic is licensed under GPL V3. It is cross-platform (Java 1.5+ required) and available at http://www.usadellab.org/cms/index.php?page=trimmomaticContact: usadel@bio1.rwth-aachen.deSupplementary information: Supplementary data are available at Bioinformatics online.

Similar Papers
  • Research Article
  • Cite Count Icon 135
  • 10.1016/j.jmoldx.2012.08.001
Detection of FLT3 Internal Tandem Duplication in Targeted, Short-Read-Length, Next-Generation Sequencing Data
  • Nov 14, 2012
  • The Journal of Molecular Diagnostics
  • David H Spencer + 8 more

Detection of FLT3 Internal Tandem Duplication in Targeted, Short-Read-Length, Next-Generation Sequencing Data

  • Research Article
  • Cite Count Icon 27
  • 10.1161/circgenetics.113.000085
Short Read (Next-Generation) Sequencing
  • Jul 14, 2013
  • Circulation: Cardiovascular Genetics
  • Jaya Punetha + 1 more

Rapid advances in DNA sequencing technologies have made it increasingly cost-effective to obtain accurate and timely large-scale genomic sequence data on individuals (short read massively parallel or next generation [next-gen]). A next-gen molecular diagnostic approach that has seen rapid deployment in the clinic over the last year is exome sequencing. Whole exome sequencing covers all protein-coding genes in the genome (≈1.1% of genome), and an exome test for a single patient generates ≈6 gigabases (109 bp) of DNA sequence data. A key challenge facing routine use of next-gen data in patient diagnosis and management is data interpretation. What sequence variant findings are relevant to diagnosis (pathogenic mutations)? What sequence variant findings are relevant to clinical care but not necessarily to patient diagnosis (clinically actionable incidental data)? What sequence information should be stored, and where can it be stored? This review provides a tutorial on current approaches to answering these questions. A recent landmark study showed that application of next-gen sequencing to a large cohort of idiopathic dilated cardiomyopathy patients found ≈27% of patients to show mutations of the titin gene, the most complex gene in the genome (363 exons). We use titin in cardiomyopathy as an exemplar for explaining next-gen sequencing approaches and data interpretation. Decreasing sequencing costs and broad dissemination of next-generation (next-gen) equipment and expertise are increasing availability of massively parallel sequencing of patient DNA samples (short read massively parallel or next-gen sequencing).1,2 Most rapidly expanding is exome sequencing, where all protein-coding sequences (exons) are selected from total genomic DNA and selectively sequenced.3 Alternative approaches to next-gen sequencing include targeted sequencing (TS) and whole genome (complete genome) sequencing. Currently, marketed targeted Sanger sequencing panels using traditional individual exon-by-exon sequencing remain expensive and time consuming, and massively parallel next-gen approaches are beginning to supplant …

  • Research Article
  • 10.1158/1538-7445.sabcs21-p2-01-15
Abstract P2-01-15: Developing highly sensitive high NGS data efficient ctDNA detection assays for breast cancer surveillance
  • Feb 15, 2022
  • Cancer Research
  • Aihua Fu + 8 more

Introduction: Growing data established the importance of monitoring dynamic changes in circulating tumor DNA (ctDNA) to identify early signs of therapeutic responses, allowing for timely management of treatment to achieve more effective personalized therapy. Higher assay accuracy and consistency, and lower assay cost will support more clinical validation trials and benefit more cancer patients with non-invasive ctDNA NGS tests that can simultaneously map multiple genomic alterations at an affordable price. Method: The NVIGEN X - Precision Cancer Profiling test is a next generation sequencing (NGS) based circulating tumor DNA detection assay using the hybridization capture approach with customized gene panels. Our ctDNA NGS assay was developed with the use of high performance magnetic nanobeads, which enhances assay workflow at key steps including cfDNA extraction, NGS library preparation, and target enrichment. Experiments with individual plasma samples and DNA mutant fragments spiked in plasma samples were carried out to establish the assay performance such as sensitivity, specificity, consistency and data efficiency. NGS data QC metrics of the NVIGEN assay were compared with other assays in peer reviewed publications. Results: We developed a focus 32 gene panel that covers 144 kb of gene regions of clinical significance for breast cancer treatment monitoring and guidance, such as AKT1, ERBB2, PIK3CA, EGFR, ESR1, BRCA1/2, and CD274. Our results demonstrated the capability of NVIGEN X ctDNA NGS assay to detect rare copies (8 cp) of gene mutation at 0.07% MAF from DNA mutant fragments spiked into plasma samples. The NVIGEN X ctDNA NGS assays consistently presented 2-5% duplication rate, >80% on-target rate, <10% CV for key NGS data metrics, and on average required 1.36X paired reads per 1X unique coverage. Compared with the Roche Avenio assays (targeted, expanded and surveillance panels) as published in 2020 which on average required 9.36X paired reads per 1X unique coverage, the NVIGEN X -precision cancer profiling assays demonstrated 85% reduction in NGS data need to generate each unique coverage. Compared with the original Capp-seq data as published in the 2014 Nature Medicine paper which required in average 13.78 or 27.56 paired reads per unique coverage, the NVIGEN X assay demonstrated >90% reduction in NGS data need per unique coverage. Conclusion: The NVIGEN X - Precision Cancer Profiling assay provided high NGS assay performance with high sensitivity, specificity, and consistency, and significantly improved NGS data efficiency. This allows for dramatically reduced assay cost and will help support routine applications of ctDNA NGS tests to improve cancer patient treatment. Experiments of applying NVIGEN X assays for clinical research with patient samples are ongoing and will be presented. Citation Format: Aihua Fu, Wenwu Cui, Minh V. Ton, Kevan Wang, Weiwei Gu, Tianhong Li, Heather A. Parsons, Minetta C. Liu, George W. Sledge. Developing highly sensitive high NGS data efficient ctDNA detection assays for breast cancer surveillance [abstract]. In: Proceedings of the 2021 San Antonio Breast Cancer Symposium; 2021 Dec 7-10; San Antonio, TX. Philadelphia (PA): AACR; Cancer Res 2022;82(4 Suppl):Abstract nr P2-01-15.

  • Research Article
  • Cite Count Icon 18
  • 10.1016/j.cllc.2022.11.010
The Performance of an Extended Next Generation Sequencing Panel Using Endobronchial Ultrasound-Guided Fine Needle Aspiration Samples in Non-Squamous Non-Small Cell Lung Cancer: A Pragmatic Study
  • Dec 5, 2022
  • Clinical Lung Cancer
  • Chenchen Zhang + 9 more

The Performance of an Extended Next Generation Sequencing Panel Using Endobronchial Ultrasound-Guided Fine Needle Aspiration Samples in Non-Squamous Non-Small Cell Lung Cancer: A Pragmatic Study

  • PDF Download Icon
  • Research Article
  • Cite Count Icon 9
  • 10.38211/joarps.2023.04.01.61
Role of Next Generation Sequencing (NGS) in Plant Disease Management: A Review
  • Feb 23, 2023
  • Journal of Applied Research in Plant Sciences
  • Muhammad Saeed + 7 more

A high throughput technique used to determine a part of the nucleotide sequence of an organism’s genome is called next generation sequencing (NGS). NGS has been Proven revolutionary in genomics. Clinical diagnostics, Plant diseases diagnostic and other aspects of medical are now made possible by sequencing. Techniques of NGS: there are different techniques of NGS which are being used in real life sciences i.e., Illumina sequencing, Pyrosequencing, Roche 454 sequencing and Ion torrent sequencing. All vintage methods like culturing in bacterial, fungal, and viral samples are being suppressed by next generation sequencing. The potential for random metagenomic sequencing of sick samples to find potential pathogens has surfaced with the development of next-generation high-throughput parallel sequencing technology. NGS enables highly efficient, rapid, low-cost DNA or RNA high-throughput sequencing of plant virus and viroids genomes, as well as specific small RNAs generated during infection. Although this technique is not so much familiar in the field of plant diseases. However, its widespread application in agronomic sciences will make it possible to create solutions to future food-related challenges that involve biotic stress.

  • Research Article
  • Cite Count Icon 46
  • 10.1016/j.jmoldx.2013.10.006
Validation for Clinical Use of, and Initial Clinical Experience with, a Novel Approach to Population-Based Carrier Screening using High-Throughput, Next-Generation DNA Sequencing
  • Dec 27, 2013
  • The Journal of Molecular Diagnostics
  • Stephanie Hallam + 12 more

Validation for Clinical Use of, and Initial Clinical Experience with, a Novel Approach to Population-Based Carrier Screening using High-Throughput, Next-Generation DNA Sequencing

  • Research Article
  • Cite Count Icon 61
  • 10.1093/bioinformatics/btr447
SEED: efficient clustering of next-generation sequences
  • Aug 2, 2011
  • Bioinformatics
  • Ergude Bao + 3 more

Motivation: Similarity clustering of next-generation sequences (NGS) is an important computational problem to study the population sizes of DNA/RNA molecules and to reduce the redundancies in NGS data. Currently, most sequence clustering algorithms are limited by their speed and scalability, and thus cannot handle data with tens of millions of reads.Results: Here, we introduce SEED—an efficient algorithm for clustering very large NGS sets. It joins sequences into clusters that can differ by up to three mismatches and three overhanging residues from their virtual center. It is based on a modified spaced seed method, called block spaced seeds. Its clustering component operates on the hash tables by first identifying virtual center sequences and then finding all their neighboring sequences that meet the similarity parameters. SEED can cluster 100 million short read sequences in <4 h with a linear time and memory performance. When using SEED as a preprocessing tool on genome/transcriptome assembly data, it was able to reduce the time and memory requirements of the Velvet/Oasis assembler for the datasets used in this study by 60–85% and 21–41%, respectively. In addition, the assemblies contained longer contigs than non-preprocessed data as indicated by 12–27% larger N50 values. Compared with other clustering tools, SEED showed the best performance in generating clusters of NGS data similar to true cluster results with a 2- to 10-fold better time performance. While most of SEED's utilities fall into the preprocessing area of NGS data, our tests also demonstrate its efficiency as stand-alone tool for discovering clusters of small RNA sequences in NGS data from unsequenced organisms.Availability: The SEED software can be downloaded for free from this site: http://manuals.bioinformatics.ucr.edu/home/seed.Contact: thomas.girke@ucr.eduSupplementary information: Supplementary data are available at Bioinformatics online

  • Research Article
  • Cite Count Icon 16
  • 10.1111/nph.13851
Data processing can mask biology: towards better reporting of fungal barcoding data?
  • Jan 28, 2016
  • New Phytologist
  • Marc‐André Selosse + 2 more

Data processing can mask biology: towards better reporting of fungal barcoding data?

  • Preprint Article
  • 10.7490/f1000research.1112971.1
Distributed file systems for storage and analysis of Next-Generation Sequencing data
  • Sep 1, 2016
  • F1000Research
  • Luca Beltrame + 4 more

Analysis of NGS (Next Generation Sequencing) data is a computationally demanding task requiring large amounts of CPU, memory, and disk space. There is also a requirement for high performance data storage systems, resilient to hardware failure, to be connected directly to the computing infrastructure (typically a multi-node cluster) to store large quantities of NGS data reliably. Traditional shared file systems such as NFS (Network File System) do not offer the performance, scalability or cache coherence required by modern NGS data analysis, so alternatives including GlusterFS, Ceph, and Lustre have been developed. However, there is a trade-off between data safety on replicated local storage and degradation of performance across distributed storage. Resilience to hardware failure is typically provided by RAID (Redundant Array of Independent Disks) and redundant storage nodes. Here we describe the evaluation of an alternative file system, RozoFS ( https://github.com/rozofs/rozofs ) for use with demanding NGS data analysis workloads. We used a synthetic data set (DREAM-TCGA data set 3) to run a complete tumor-normal analysis pipeline (“bcbio”, https://github.com/chapmanb/bcbio-nextgen ), including base quality recalibration, local indel realignment, somatic variant calling, and structural variants as a benchmark to compare RozoFS with a traditional shared file system (NFS) on two different HPC (High Performance Computing; Cloud4CaRE project) clusters. Our results show high reliability and good performance of RozoFS compared to NFS, in particular, during heavy I/O workloads. These findings indicate that the reliability and robustness of RozoFS's make it a good candidate for demanding NGS analysis workloads.

  • Research Article
  • Cite Count Icon 117
  • 10.1111/j.1469-8137.2011.03755.x
Towards standardization of the description and publication of next‐generation sequencing datasets of fungal communities
  • May 9, 2011
  • New Phytologist
  • R Henrik Nilsson + 12 more

Towards standardization of the description and publication of next‐generation sequencing datasets of fungal communities

  • Research Article
  • Cite Count Icon 3
  • 10.1371/journal.pdig.0000825
Addressing data management and analysis challenges in viral genomics: The Swiss HIV cohort study viral next generation sequencing database.
  • Apr 21, 2025
  • PLOS digital health
  • Marius Zeeb + 14 more

Numerous HIV related outcomes can be determined on the viral genome, for example, resistance associated mutations, population transmission dynamics, viral heritability traits, or time since infection. Viral sequences of people with HIV (PWH) are therefore essential for therapeutic and research purposes. While in the first three decades of the HIV pandemic viral genomes were mainly sequenced using Sanger sequencing, the last decade has seen a shift towards next-generation sequencing (NGS) as the preferred method. NGS can achieve near full length genome sequence coverage and simultaneously, it accurately encapsulates the within-host diversity by characterizing HIV subpopulations. NGS opens new avenues for HIV research, but it also presents challenges concerning data management and analysis. We therefore set up the Swiss HIV Cohort Study Viral NGS Database (SHCND) to address key issues in the handling of NGS data including high loads of raw- and processed NGS data, data storage solutions, downstream application of sophisticated bioinformatic tools, high-performance computing resources, and reproducibility. The database is nested within the Swiss HIV Cohort Study (SHCS) and the Zurich Primary HIV Infection Cohort Study (ZPHI), which together enrolled 21,876 PWH since 1988 and include a biobank dating back to the early nineties. Since its initiation in 2018, the SHCND accumulated NGS sequences (plasma and proviral origin) of 5,178 unique PWH. We here describe the design, set-up, and use of this NGS database. Overall, the SHCND has contributed to several research projects on HIV pathogenesis, treatment, drug resistance, and molecular epidemiology, and has thereby become a central part of HIV-genomics research in Switzerland.

  • Research Article
  • Cite Count Icon 67
  • 10.1016/j.mimet.2013.07.002
E-probe Diagnostic Nucleic acid Analysis (EDNA): A theoretical approach for handling of next generation sequencing data for diagnostics
  • Jul 16, 2013
  • Journal of Microbiological Methods
  • Anthony H Stobbe + 8 more

E-probe Diagnostic Nucleic acid Analysis (EDNA): A theoretical approach for handling of next generation sequencing data for diagnostics

  • PDF Download Icon
  • Research Article
  • Cite Count Icon 33
  • 10.1186/s12859-019-2936-9
FastProNGS: fast preprocessing of next-generation sequencing reads
  • Jun 17, 2019
  • BMC Bioinformatics
  • Xiaoshuang Liu + 5 more

BackgroundNext-generation sequencing technology is developing rapidly and the vast amount of data that is generated needs to be preprocessed for downstream analyses. However, until now, software that can efficiently make all the quality assessments and filtration of raw data is still lacking.ResultsWe developed FastProNGS to integrate the quality control process with automatic adapter removal. Parallel processing was implemented to speed up the process by allocating multiple threads. Compared with similar up-to-date preprocessing tools, FastProNGS is by far the fastest. Read information before and after filtration can be output in plain-text, JSON, or HTML formats with user-friendly visualization.ConclusionsFastProNGS is a rapid, standardized, and user-friendly tool for preprocessing next-generation sequencing data within minutes. It is an all-in-one software that is convenient for bulk data analysis. It is also very flexible and can implement different functions using different user-set parameter combinations.

  • Research Article
  • 10.4233/uuid:1752b8ce-631b-4127-91c9-92538e34a13b
Accelerating DNA Variant Calling Algorithms on High Performance Computing Systems
  • Dec 17, 2018
  • Research Repository (Delft University of Technology)
  • Siying Ren

Next generation sequencing (NGS) technologies have transformed the landscape of genomic research. With the significant advances in NGS technologies, DNA sequencing is more affordable and accessible than ever before. Meanwhile, many DNA sequence analysis tools have been developed to derive useful information from the raw sequencing data produced by NGS platforms. However, the massive amount of generated sequencing data poses a great computational challenge, thereby shifting the bottleneck towards the efficiency of the DNA sequence analysis tools. Due to the high computational needs, high performance systems are playing an important role for DNA sequence analysis. Moreover, dedicated hardware, including graphics processing units (GPUs) and field programmable gate arrays (FPGAs), have become important computational resources in many high performance systems. In this thesis, we use GPUs and FPGAs to accelerate a number of important bioinformatics algorithms. These represent the most computationally intensive algorithms of the GATK HaplotypeCaller (HC), which we use to improve its performance. GATK HC is a widely used DNA sequence analysis tool. By investigating GATK HC, three computationally intensive algorithms are selected, including the de Buijn graph (DBG) construction algorithm for micro-assembly, the pair-HMMs forward algorithm and the semi-global pairwise alignment algorithm. We first propose a novel GPU-based implementation of the DBG construction algorithm for micro-assembly. Compared with the software-only implementation, it achieves a speedup of up to 3x using synthetic datasets and a speedup of up to 2.66x using human genome datasets. We then propose a systolic array design to accelerate the pair-HMMs forward algorithm on FPGAs. Experimental results show that the FPGA-based implementation is up to 67x faster than the software-only implementation. In order to fully utilize the computing resources on FPGAs, we present a model to describe the performance characteristics of the systolic array design. Based on the analysis, we propose a novel architecture to better utilize the computing resources on FPGAs. The implementation achieves up to 90\% of the theoretical throughput for a real dataset. Next, we propose several GPU-based implementations of the pair-HMMs forward algorithm. Experimental results show that the GPU-based implementations of the pair-HMMs forward algorithm achieve a speedup of up to 5.47x over existing GPU-based implementations. Finally, we propose to accelerate the semi-global pairwise sequence alignment algorithm with traceback to obtain the optimal alignment on GPUs. Experimental results show that the GPU-based implementation is up to 14.14x faster than the software-only implementation. After accelerating these algorithms on GPUs and FPGAs, we integrate two GPU-based implementations into GATK HC. We first integrate the GPU-based implementation of the pair-HMMs forward algorithm into GATK HC. In single-threaded mode, the GPU-based GATK HC implementation is 1.71x faster than the baseline GATK HC implementation. For multi-process mode, a load-balanced multi-process optimization is proposed to ensure a more equal distribution of computation load between different processes. The GPU-based GATK HC implementation achieves up to 2.04x in load-balanced multi-process mode over the baseline GATK HC implementation in non-load-balanced multi-process mode. Next, we additionally integrated the GPU-based implementation of the semi-global alignment algorithm into the GATK HC. Experimental results shown that this implementation is 2.3x faster than the baseline GATK HC implementation in single-thread mode.

  • Research Article
  • 10.1097/01.hs9.0000562332.26969.1f
PS1009 INDIVIDUALIZED FOLLOW‐UP IN ADULT ACUTE MYELOID LEUKEMIA USING NGS
  • Jun 1, 2019
  • HemaSphere
  • Z Blasco Iturri + 17 more

Background:The current standard for morphologic complete remission in acute myeloid leukemia (AML) is less than 5% myeloblasts, but mounting data show this criterion is not sufficiently sound. Alternative methods, such as quantitative reverse‐transcription polymerase chain reaction (RT‐qPCR), are widely used to detect molecular responses, but it relies on the initial detection of a fusion transcript, or overexpressed gene. Deeper knowledge of the clonal dynamics of AML could potentially be of clinical utility. A wide scope testing technology is required in order to address the molecular heterogeneity of AML. We reasoned that an appropriate Next Generation Sequencing (NGS) panel could be a useful tool to provide personalized molecular monitoring in patients diagnosed as or progress to AML.Aims:The aim of this study is to evaluate the clinical utility of NGS panel in the prognostic and treatment monitoring in patients diagnosed as or progress to AML.Methods:We studied the genomic alterations of 19 AML cases (13 de novo, and 6 secondary to a preexisting MN) during disease follow‐up; 11 of these patients received hematopoietic stem cell transplantation (HSCT). The 67 samples were tested with our custom Pan‐Myeloid Panel (48 genes, SOPHiA GENETICS). Samples were provided by the Biobank of the University of Navarra and were processed following SOP approved by the Ethical and Scientific Committee of the University. Libraries were pair‐end sequenced on a Miseq sequencer (Illumina). Sequencing data were analyzed by two geneticists with expertise in hematological malignancies.Results:Sequencing data identified genomic clonal markers with clinical utility (i.e. diagnostic, prognostic, and/or predictive value) in 89,5% of cases. In patients not receiving HSCT (n = 8), NGS was useful to classify them in two genetic profiles: those achieving molecular complete remission (mCR) (n = 2) (Figure 1A), and those not responding to treatment and undergoing disease progression (n = 6) (Figure 1B). In the last group, NGS identified pathogenic variants in DDX41, DNMT3A, IDH1, JAK2, NRAS, SRSF2, U2AF1 genes. In patients receiving HSCT (n = 11), NGS also classified patients in two groups: those clearing pathogenic variants upon HSCT (n = 5) (Figure 1C), and those with persisting variants, not achieving mCR (n = 6) (Figure 1D). Again, NGS identified initial clones harboring pathogenic variants, like KRAS, that appeared in the 66% of the patients after HSCT failure. Also mutations in CBL, DNMT3A, FLT3, JAK2, KRAS and SRSF2 genes are present in this cohort of patients.In two cases, NGS either did not detect any clinically relevant variant, or it detected variants only after disease progression; in these two cases an NGS panel was insufficient, and therefore more comprehensive studies are needed (e.g. exomes). Of note, NGS data detected clones harboring pathogenic variants in two patients with negative minimal residual disease (MRD), as measured by flow cytometry (Figure 1D), indicating that NGS could complement current gold standard follow‐up method in some instances.Summary/Conclusion:A 48‐gene panel NGS has been useful for molecular diagnosis, treatment follow‐up, and relapse detection in nearly 90% of the AML cases included in our study (17 of 19). NGS was also useful for following mutational clearance and/or clonal evolution in 12 of 19 patients (63%), including cases undergoing HSCT. According to our data, NGS could be of clinical utility for routine diagnosis and follow‐up in an elevated proportion of AML patients, even complementing immunophenotypic techniques for MRD monitoring in some instances.image

Save Icon
Up Arrow
Open/Close
Notes

Save Important notes in documents

Highlight text to save as a note, or write notes directly

You can also access these Documents in Paperpal, our AI writing tool

Powered by our AI Writing Assistant