PhiSpy: a novel algorithm for finding prophages in bacterial genomes that combines similarity- and composition-based strategies
Prophages are phages in lysogeny that are integrated into, and replicated as part of, the host bacterial genome. These mobile elements can have tremendous impact on their bacterial hosts’ genomes and phenotypes, which may lead to strain emergence and diversification, increased virulence or antibiotic resistance. However, finding prophages in microbial genomes remains a problem with no definitive solution. The majority of existing tools rely on detecting genomic regions enriched in protein-coding genes with known phage homologs, which hinders the de novo discovery of phage regions. In this study, a weighted phage detection algorithm, PhiSpy was developed based on seven distinctive characteristics of prophages, i.e. protein length, transcription strand directionality, customized AT and GC skew, the abundance of unique phage words, phage insertion points and the similarity of phage proteins. The first five characteristics are capable of identifying prophages without any sequence similarity with known phage genes. PhiSpy locates prophages by ranking genomic regions enriched in distinctive phage traits, which leads to the successful prediction of 94% of prophages in 50 complete bacterial genomes with a 6% false-negative rate and a 0.66% false-positive rate.
- Conference Article
- 10.18699/bgrs2024-8.4-20
- Sep 19, 2024
Motivation and Aim: CRISPR-cas systems are incredibly diverse and currently are classified into six major types and over 30 subtypes [1].Apart from their role in adaptive immunity, some of the CRISPR-cas subtypes are also involved in host gene regulation and collateral damage, leading to bacteriostatic or lethal outcomes for the host.CRISPR array spacers direct and influence canonical and non-canonical functions of the CRISPR-Cas system together with subtype Cas proteins.A better understanding of spacer adaptation mechanisms is crucial for uncovering the intricacies of an evolutionary arms race between prokaryotes and phages.Methods and Algorithms: Bacterial and viral genomes were retrieved using NCBI Datasets API [2].CRISPRidentify and CRISPRcasIdentifier tools were used for CRISPR array, Cas genes detection, and subtyping [3][4].Viral genomes were mapped to their hosts using the latest version of the Virus-Host DB [5].Mapping was performed on the genus level of the hosts' phylogenetic trees.Gumbel extreme value distribution was used to determine the statistical significance of each spacer Smith-Waterman alignment score.Mobile genetic elements (MGE) and prophage predictions were obtained for all complete bacterial genomes using the VRprofile2 tool [6].Results: We present a large-scale analysis of CRISPR array spacers from 31,845 complete bacterial genomes.Differences in melting energy and GC content between identified spacers, origin bacterial genomes, and infecting bacteriophages were explored for different CRISPR-cas subtypes and bacterial genera.Significant GC content differences can be observed between bacterial genomes and their spacers and between spacers and infecting phage genomes.Spacers from the extremes of the GC content distribution were aligned to bacterial and infecting phage genomes to determine their origin.We obtain that high GC content spacers have better alignments to phage genomes than low GC spacers and that high and low GC spacers show better alignments to origin bacterial genomes than infecting phage genomes.We further tested if spacer alignment locations to origin bacterial genomes correspond to predicted MGE and Prophage locations and if alignment scores between those spacer groups differ.We found that low GC spacers do not preferentially target integrated prophage or MGE locations over the rest of the bacterial genomes, while high GC spacers tend to target integrated prophage sequences over MGE and remaining bacterial genome sequences. Conclusion:The GC content of the spacers was smaller than the GC content of the source bacterial genome but larger than the infecting viral genome.This observation aligns with the hypothesis that the majority of CRISPR spacers were adapted from the bacteriophage genomes and serve a canonical function.Alignments of the spacers from GC-rich distribution tails have shown their preferential targeting of host genomes.This further supports the hypothesis that GC-rich spacers originated from the bacterial genome and BGRS/SB-2024
- Research Article
127
- 10.1093/nar/gkac321
- May 7, 2022
- Nucleic Acids Research
VRprofile2 is an updated pipeline that rapidly identifies diverse mobile genetic elements in bacterial genome sequences. Compared with the previous version, three major improvements were made. First, the user-friendly visualization could aid users in investigating the antibiotic resistance gene cassettes in conjunction with various mobile elements in the multiple resistance region with mosaic structure. VRprofile2 could compare the predicted mobile elements to the collected known mobile elements with similar architecture. A new mobilome indicator was proposed to give an overall estimation of the mobilome size in individual bacterial genomes. Second, the relationship between antibiotic resistance genes, mobile elements, and host strains would be efficiently examined with the aid of predicted strain's sequence typing, the incompatibility group and the transferability of plasmids. Finally, the updated back-end database, MobilomeDB2, now collected nearly a thousand active mobile elements retrieved from literature or based on prediction. The pre-computed results of the antibiotic resistance gene-carrying mobile elements of >5500 ESKAPEE genomes were also provided. We expect that VRprofile2 will provide better support for researchers interested in bacterial mobile elements and the dissemination of antibiotic resistance. VRprofile2 is freely available to all users without any login requirement at https://tool2-mml.sjtu.edu.cn/VRprofile.
- Research Article
321
- 10.1038/s41467-021-22757-1
- Apr 23, 2021
- Nature Communications
Antibiotic resistance spreads among bacteria through horizontal transfer of antibiotic resistance genes (ARGs). Here, we set out to determine predictive features of ARG transfer among bacterial clades. We use a statistical framework to identify putative horizontally transferred ARGs and the groups of bacteria that disseminate them. We identify 152 gene exchange networks containing 22,963 bacterial genomes. Analysis of ARG-surrounding sequences identify genes encoding putative mobilisation elements such as transposases and integrases that may be involved in gene transfer between genomes. Certain ARGs appear to be frequently mobilised by different mobile genetic elements. We characterise the phylogenetic reach of these mobilisation elements to predict the potential future dissemination of known ARGs. Using a separate database with 472,798 genomes from Streptococcaceae, Staphylococcaceae and Enterobacteriaceae, we confirm 34 of 94 predicted mobilisations. We explore transfer barriers beyond mobilisation and show experimentally that physiological constraints of the host can explain why specific genes are largely confined to Gram-negative bacteria although their mobile elements support dissemination to Gram-positive bacteria. Our approach may potentially enable better risk assessment of future resistance gene dissemination.
- Research Article
47
- 10.1371/journal.pgen.1010065
- Feb 14, 2022
- PLOS Genetics
Most bacterial genomes contain horizontally acquired and transmissible mobile genetic elements, including temperate bacteriophages and integrative and conjugative elements. Little is known about how these elements interact and co-evolved as parts of their host genomes. In many cases, it is not known what advantages, if any, these elements provide to their bacterial hosts. Most strains of Bacillus subtilis contain the temperate phage SPß and the integrative and conjugative element ICEBs1. Here we show that the presence of ICEBs1 in cells protects populations of B. subtilis from predation by SPß, likely providing selective pressure for the maintenance of ICEBs1 in B. subtilis. A single gene in ICEBs1 (yddK, now called spbK for SPß killing) was both necessary and sufficient for this protection. spbK inhibited production of SPß, during both activation of a lysogen and following de novo infection. We found that expression spbK, together with the SPß gene yonE constitutes an abortive infection system that leads to cell death. spbK encodes a TIR (Toll-interleukin-1 receptor)-domain protein with similarity to some plant antiviral proteins and animal innate immune signaling proteins. We postulate that many uncharacterized cargo genes in ICEs may confer selective advantage to cells by protecting against other mobile elements.
- Research Article
71
- 10.3389/fmicb.2023.1179966
- May 15, 2023
- Frontiers in Microbiology
Genome-based analysis is crucial in monitoring antibiotic-resistant bacteria (ARB)and antibiotic-resistance genes (ARGs). Short-read sequencing is typically used to obtain incomplete draft genomes, while long-read sequencing can obtain genomes of multidrug resistance (MDR) plasmids and track the transmission of plasmid-borne antimicrobial resistance genes in bacteria. However, long-read sequencing suffers from low-accuracy base calling, and short-read sequencing is often required to improve genome accuracy. This increases costs and turnaround time. In this study, a novel ONT sequencing method is described, which uses the latest ONT chemistry with improved accuracy to assemble genomes of MDR strains and plasmids from long-read sequencing data only. Three strains of Salmonella carrying MDR plasmids were sequenced using the ONT SQK-LSK114 kit with flow cell R10.4.1, and de novo genome assembly was performed with average read accuracy (Q > 10) of 98.9%. For a 5-Mb-long bacterial genome, finished genome sequences with accuracy of >99.99% could be obtained at 75× sequencing coverage depth using Flye and Medaka software. Thus, this new ONT method greatly improves base-calling accuracy, allowing for the de novo assembly of high-quality finished bacterial or plasmid genomes without the need for short-read sequencing. This saves both money and time and supports the application of ONT data in critical genome-based epidemiological analyses. The novel ONT approach described in this study can take the place of traditional combination genome assembly based on short- and long-read sequencing, enabling pangenomic analyses based on high-quality complete bacterial and plasmid genomes to monitor the spread of antibiotic-resistant bacteria and antibiotic resistance genes.
- Research Article
- 10.2174/0122115501397209251118052633
- Nov 22, 2025
- Current Biotechnology
Introduction: The increasing emergence of zoonotic pathogens and antimicrobial resistance (AMR) highlights the need for rapid and accurate computational tools to assess the zoonotic potential of bacterial strains. In this study, we present Zoonomix, a bioinformatics pipeline designed to detect and rank genes associated with pathogenicity, virulence, and antibiotic resistance, thereby enabling risk assessment for zoonotic transmission. Methods: Zoonomix integrates a curated database of ~25,000 genes related to adherence, biofilm formation, efflux pumps, exotoxins, resistance, integrative and conjugative elements (ICEs), and secretion systems (T3SS, T4SS, and T6SS). It uses BLASTN and a scoring algorithm to assess pathogenicity and HGT risk, classifying bacterial strains into low, moderate, or high risk, with insights into antibiotic resistance migration. Results: When analyzing 60 whole genome sequences of both zoonotic and non-zoonotic bacterial species using the Zoonomix pipeline, over 90% of the results were accurately classified in accordance with existing literature. Notably, the pipeline predicted a potential future zoonotic and pathogenic capability for bacterial species such as A. pleuropneumoniae and M. haemolytica. Discussion: Zoonomix offers a comprehensive framework for assessing zoonotic potential and antibiotic resistance by integrating genomics, bioinformatics, and predictive analytics. Its ability to detect current gene status, identify mutation-prone genes, summarize mutation hotspots, and flag horizontal gene transfer events make it a valuable tool for disease surveillance and outbreak prevention. conclusion: Zoonomix is a scalable, open-source tool for assessing zoonotic potential and AMR risk in bacterial genomes. By detecting key genes, predicting future mutations, and flagging ICE-mediated resistance transfer, it offers valuable insights for genomic epidemiology and public health surveillance. The open-source pipeline is available at https://github.com/Umeshkumarku1/ZoonomiX. Conclusion: Zoonomix is a scalable, open-source bioinformatics pipeline designed to assess the zoonotic potential and antimicrobial resistance (AMR) risk in bacterial whole genome sequences. By detecting key genes associated with zoonosis, identifying markers that predict the future pathogenic or zoonotic potential of bacteria, and flagging integrative conjugative element (ICEs)- mediated resistance gene transfer mechanisms, the tool provides comprehensive insights into bacterial threats. The pipeline's source code and documentation are freely available for the research community at the following GitHub repository: https://github.com/Umeshkumarku1/ZoonomiX.
- Research Article
16
- 10.1038/sj.embor.embor930
- Sep 1, 2003
- EMBO reports
After more than 150 years of research in microbiology, new technologies and new insights into the microbial world have sparked a revolution in the field. This is a much needed development, not only to renew interest in prokaryote research, but also to meet many emerging challenges in medicine, agriculture and industrial processes. Although many microbiologists—such as Emil von Behring, Robert Koch, Jacques Monod, Francois Jacob, Andre Lwoff, Alexander Fleming, Selman A. Waksman and Joshua Lederberg—grace the list of Nobel laureates, attention moved away from microbiology as biologists focused their interest on eukaryotic cells and higher organisms in the 1970s and 1980s. Furthermore, from the beginning, research on prokaryotes has suffered from an anthropocentric view, regarding as interesting only those organisms that cause disease or that can be exploited for industrial or agricultural use. But the advent of new technologies, some of which have been driven by a need to understand eukaryotes, may change this. We are increasingly realizing how little we know about microbes in general, their diversity, the mechanisms of their evolution and adaptation and their modes of existence within, and communication with, their environment and higher organisms. As bacteria have succeeded in occupying virtually all ecological niches on this planet, ranging from arctic regions to oceanic hot springs, they hold an immense wealth of genetic information that we have barely started to explore and that may provide many useful applications for humans. The new technologies that allow us to sequence and annotate whole genomes more rapidly and to analyse the expression of thousands of genes in a single experiment are likely to speed up this change, particularly as microbes are well suited for high‐throughput analysis. Any microbial genome can now be sequenced within a few hours and, in the near future, new bioinformatics tools will enable scientists not …
- Research Article
14
- 10.3389/fmicb.2016.00469
- Apr 12, 2016
- Frontiers in Microbiology
Several metagenomic projects have been accomplished or are in progress. However, in most cases, it is not feasible to generate complete genomic assemblies of species from the metagenomic sequencing of a complex environment. Only a few studies have reported the reconstruction of bacterial genomes from complex metagenomes. In this work, Binning-Assembly approach has been proposed and demonstrated for the reconstruction of bacterial and viral genomes from 72 human gut metagenomic datasets. A total 1156 bacterial genomes belonging to 219 bacterial families and, 279 viral genomes belonging to 84 viral families could be identified. More than 80% complete draft genome sequences could be reconstructed for a total of 126 bacterial and 11 viral genomes. Selected draft assembled genomes could be validated with 99.8% accuracy using their ORFs. The study provides useful information on the assembly expected for a species given its number of reads and abundance. This approach along with spiking was also demonstrated to be useful in improving the draft assembly of a bacterial genome. The Binning-Assembly approach can be successfully used to reconstruct bacterial and viral genomes from multiple metagenomic datasets obtained from similar environments.
- Research Article
- 10.22067/jpp.v29i4.45988
- Jan 12, 2016
- SHILAP Revista de lepidopterología
توانمندی تشخیص و ردیابی صحیح فیتوپلاسماها از گیاهان، اهمیت به سزایی در کنترل موثر این عوامل و جلوگیری از گسترش آنها از طریق مواد گیاهی آلوده دارد. استخراج DNA به گونه ای که کمترین مواد ممانعت کننده و بیشترین میزان ماده ژنتیکی فیتوپلاسما حاصل گردد، در موفقیت آزمون واکنش زنجیره ای پلیمراز و صحت نتایج حاصل از آن تاثیرگذار است. در این تحقیق، مقایسه روش های مختلف استخراج DNA نشان داد که استفاده از روش مناسب، در کاهش واکنش منفی کاذب اهمیت دارد. اما حتی در روش موفق استخراج DNA مبتنی بر استفاده از ستون نیز واکنش های منفی کاذب هرچند به تعداد اندک اما به صورت پراکنده وجود دارد. این امر می تواند ناشی از غلظت پایین و پراکنش نامنظم سلول-های فیتوپلاسمایی در بافتهای گیاه میزبان باشد. بر اساس نتایج حاصل از تعیین توالی قطعه تکثیر شده با ترکیب آغازگری P1/P7-R16F2n/R16R2، واکنش مثبت دروغین در برخی تکرارهای آزمون Nested PCR به دست آمد. به منظور کاهش آلودگی های متقاطع و اجتناب از واکنش های مثبت دروغین، آزمون Single Tube Nested PCR بهینه سازی گردید. علیرغم دستیابی به محاسن مختلف این روش از جمله سهولت اجرا، صرفه جویی در هزینه و زمان و همچنین توان ردیابی بیمارگر در غلظت های پایین، متاسفانه واکنش های مثبت کاذب در تعداد کمی از نمونه ها همچنان مشاهده شد. در مجموع نتایج حاصل از روش Nested PCR با یک ترکیب آغازگری و به تنهایی برای ردیابی و ارزیابی-های وسیع باغات و مزارع کفایت نمی کند و استفاده از ترکیبات پرایمری مختلف، تعیین توالی و یا هضم آنزیمی به منظور اجتناب از واکنش مثبت دروغین توصیه می گردد.
- Research Article
5
- 10.1016/j.immuni.2018.08.023
- Sep 1, 2018
- Immunity
Viral Anti-CRISPR Tactics: No Success without Sacrifice
- Research Article
52
- 10.1186/1471-2105-6-171
- Jul 12, 2005
- BMC Bioinformatics
BackgroundPublic databases now contain multitude of complete bacterial genomes, including several genomes of the same species. The available data offers new opportunities to address questions about bacterial genome evolution, a task that requires reliable fine comparison data of closely related genomes. Recent analyses have shown, using pairwise whole genome alignments, that it is possible to segment bacterial genomes into a common conserved backbone and strain-specific sequences called loops.ResultsHere, we generalize this approach and propose a strategy that allows systematic and non-biased genome segmentation based on multiple genome alignments. Segmentation analyses, as applied to 13 different bacterial species, confirmed the feasibility of our approach to discern the 'mosaic' organization of bacterial genomes. Segmentation results are available through a Web interface permitting functional analysis, extraction and visualization of the backbone/loops structure of documented genomes. To illustrate the potential of this approach, we performed a precise analysis of the mosaic organization of three E. coli strains and functional characterization of the loops.ConclusionThe segmentation results including the backbone/loops structure of 13 bacterial species genomes are new and available for use by the scientific community at the URL: .
- Supplementary Content
79
- 10.1111/mmi.14167
- Dec 25, 2018
- Molecular Microbiology
SummaryThanks to the exponentially increasing number of publicly available bacterial genome sequences, one can now estimate the important contribution of integrated viral sequences to the diversity of bacterial genomes. Indeed, temperate bacteriophages are able to stably integrate the genome of their host through site‐specific recombination and transmit vertically to the host siblings. Lysogenic conversion has been long acknowledged to provide additional functions to the host, and particularly to bacterial pathogen genomes where prophages contribute important virulence factors. This review aims particularly at highlighting the current knowledge and questions about lysogeny in Salmonella genomes where functional prophages are abundant, and where genetic interactions between host and prophages are of particular importance for human health considerations.
- Research Article
27
- 10.1016/j.bbapap.2013.12.008
- Dec 22, 2013
- Biochimica et Biophysica Acta (BBA) - Proteins and Proteomics
Anti-restriction and anti-modification (anti-RM) is the ability to prevent cleavage by DNA restriction–modification (RM) systems of foreign DNA entering a new bacterial host. The evolutionary consequence of anti-RM is the enhanced dissemination of mobile genetic elements. Homologues of ArdA anti-RM proteins are encoded by genes present in many mobile genetic elements such as conjugative plasmids and transposons within bacterial genomes. The ArdA proteins cause anti-RM by mimicking the DNA structure bound by Type I RM enzymes. We have investigated ArdA proteins from the genomes of Enterococcus faecalis V583, Staphylococcus aureus Mu50 and Bacteroides fragilis NCTC 9343, and compared them to the ArdA protein expressed by the conjugative transposon Tn916. We find that despite having very different structural stability and secondary structure content, they can all bind to the EcoKI methyltransferase, a core component of the EcoKI Type I RM system. This finding indicates that the less structured ArdA proteins become fully folded upon binding. The ability of ArdA from diverse mobile elements to inhibit Type I RM systems from other bacteria suggests that they are an advantage for transfer not only between closely-related bacteria but also between more distantly related bacterial species.
- Research Article
523
- 10.1093/nar/29.18.3742
- Sep 15, 2001
- Nucleic Acids Research
Restriction-modification (RM) systems are composed of genes that encode a restriction enzyme and a modification methylase. RM systems sometimes behave as discrete units of life, like viruses and transposons. RM complexes attack invading DNA that has not been properly modified and thus may serve as a tool of defense for bacterial cells. However, any threat to their maintenance, such as a challenge by a competing genetic element (an incompatible plasmid or an allelic homologous stretch of DNA, for example) can lead to cell death through restriction breakage in the genome. This post-segregational or post-disturbance cell killing may provide the RM complexes (and any DNA linked with them) with a competitive advantage. There is evidence that they have undergone extensive horizontal transfer between genomes, as inferred from their sequence homology, codon usage bias and GC content difference. They are often linked with mobile genetic elements such as plasmids, viruses, transposons and integrons. The comparison of closely related bacterial genomes also suggests that, at times, RM genes themselves behave as mobile elements and cause genome rearrangements. Indeed some bacterial genomes that survived post-disturbance attack by an RM gene complex in the laboratory have experienced genome rearrangements. The avoidance of some restriction sites by bacterial genomes may result from selection by past restriction attacks. Both bacteriophages and bacteria also appear to use homologous recombination to cope with the selfish behavior of RM systems. RM systems compete with each other in several ways. One is competition for recognition sequences in post-segregational killing. Another is super-infection exclusion, that is, the killing of the cell carrying an RM system when it is infected with another RM system of the same regulatory specificity but of a different sequence specificity. The capacity of RM systems to act as selfish, mobile genetic elements may underlie the structure and function of RM enzymes.
- Book Chapter
- 10.1007/978-981-19-5224-1_68
- Nov 6, 2022
This paper describes the design, development, and implementation of a standalone bioinformatic tool for the prediction of putative prophage loci in bacterial host genomes using statistical measures, based on the algorithm published as the “Prophage Loci Predictor for Bacterial Genomes” and described as the loci predictor algorithm. This algorithm proposed a novel approach to the problem of detecting prophage regions in bacterial genomic information using particle swarm optimization, using a fixed size pattern lookup table to detect virus-like pattern distributions in the host/bacterial genome. As this algorithm was designed with the intension of providing highly consistent and fast performance, the time-to-process sequence is the primary metric for evaluating the performance of the tool, and the processing speed was expected to scale only with the size of the genome under consideration and not on the size of the pattern database as is the case with other algorithms in its class. The implemented tool was evaluated using both the algorithms test and training sets and was shown to obtain a linear co-related performance as expected in both training and prediction phases of the performance testing.KeywordsBioinformaticsGenomicsProphagesPSO