DeePaC: predicting pathogenic potential of novel DNA with reverse-complement neural networks.

Jakub M Bartoszewicz,Robert Rentzsch,Anja Seidel,Bernhard Y Renard,Inanc Birol

doi:10.1093/bioinformatics/btz541

Jakub M Bartoszewicz, Robert Rentzsch + Show 3 more

Open Access

https://doi.org/10.1093/bioinformatics/btz541

Copy DOI

Journal: Bioinformatics	Publication Date: Jul 12, 2019
Citations: 44	License type: cc-by-nd

Affiliation: Robert Koch Institute, Freie Universität Berlin

Abstract

We expect novel pathogens to arise due to their fast-paced evolution, and new species to be discovered thanks to advances in DNA sequencing and metagenomics. Moreover, recent developments in synthetic biology raise concerns that some strains of bacteria could be modified for malicious purposes. Traditional approaches to open-view pathogen detection depend on databases of known organisms, which limits their performance on unknown, unrecognized and unmapped sequences. In contrast, machine learning methods can infer pathogenic phenotypes from single NGS reads, even though the biological context is unavailable. We present DeePaC, a Deep Learning Approach to Pathogenicity Classification. It includes a flexible framework allowing easy evaluation of neural architectures with reverse-complement parameter sharing. We show that convolutional neural networks and LSTMs outperform the state-of-the-art based on both sequence homology and machine learning. Combining a deep learning approach with integrating the predictions for both mates in a read pair results in cutting the error rate almost in half in comparison to the previous state-of-the-art. The code and the models are available at: https://gitlab.com/rki_bioinformatics/DeePaC. Supplementary data are available at Bioinformatics online.

Full Text