InParanoid 8: orthology analysis between 273 proteomes, mostly eukaryotic.
The InParanoid database (http://InParanoid.sbc.su.se) provides a user interface to orthologs inferred by the InParanoid algorithm. As there are now international efforts to curate and standardize complete proteomes, we have switched to using these resources rather than gathering and curating the proteomes ourselves. InParanoid release 8 is based on the 66 reference proteomes that the ‘Quest for Orthologs’ community has agreed on using, plus 207 additional proteomes from the UniProt complete proteomes—in total 273 species. These represent 246 eukaryotes, 20 bacteria and seven archaea. Compared to the previous release, this increases the number of species by 173% and the number of pairwise species comparisons by 650%. In turn, the number of ortholog groups has increased by 423%. We present the contents and usages of InParanoid 8, and a detailed analysis of how the proteome content has changed since the previous release.
- Research Article
- 10.1007/s00705-022-05402-0
- Mar 5, 2022
- Archives of virology
Analysis of orthology is important for understanding protein conservation, function, and phylogenomics. In this study, we performed a comprehensive analysis of gene orthology in the family Ascoviridae based on identification of 366 protein homologue groups and phylogenetic analysis of 34 non-single-copy proteins. Our findings revealed 90 newly annotated proteins, five newly identified core proteins for the family Ascoviridae, and 14 core proteins for the genus Ascovirus. A phylogenomic tree of 11 Ascoviridae members was constructed based on a concatenation of 35 of the 45 ortholog groups. In combination with phosphoproteomic results and conservation estimations, 30 conserved phosphorylation sites on 17 phosphoproteins were identified from a total of 176 phosphosites on 57 phosphoproteins from Heliothis virescens ascovirus 3h (HvAV-3h), providing potential research targets for investigating the role of these protein in the regulation of viral infection. This study will facilitate genome annotation and comparison of further Ascoviridae members as well as functional genomic investigations.
- Research Article
13
- 10.1142/s0219720012500205
- Oct 18, 2012
- Journal of Bioinformatics and Computational Biology
Reference proteomes are generated by increasingly sophisticated annotation pipelines as part of regular genome build releases; yet, the corresponding changes in reference proteomes' content are dramatic. In the history of the NCBI-curated human proteome, the total number of entries has remained roughly constant but approximately half of the proteins from the 2003 build 33 are no longer represented by entries in current releases, while about the same number of new proteins have been added (for sequence identity thresholds 50-90%). Although mostly hypothetical proteins are affected, there are also spectacular cases of entry removal/addition of well studied proteins. The changes between the 2003 and recent human proteomes are in a similar order of magnitude as the differences between recent human and chimpanzee proteome releases. As an application example, we show that the proteome fluctuations affect the interpretation (about 74% of hits) of organelle-specific mass-spectrometry data. Although proteome quality tends to improve with more recent releases as, for example, the fraction of proteins with functional annotation has increased over time, existing evidence implies that, apparently, the proteome content still remains incomplete, not just pertaining to isoforms/sequence variants but also to proteins and their families that are clearly distinct.
- Research Article
10
- 10.1016/j.tcs.2017.06.017
- Jun 27, 2017
- Theoretical Computer Science
A new distributed alignment-free approach to compare whole proteomes
- Research Article
35
- 10.1093/gigascience/giz118
- Oct 1, 2019
- GigaScience
BackgroundGene homology type classification is required for many types of genome analyses, including comparative genomics, phylogenetics, and protein function annotation. Consequently, a large variety of tools have been developed to perform homology classification across genomes of different species. However, when applied to large genomic data sets, these tools require high memory and CPU usage, typically available only in computational clusters.FindingsHere we present a new graph-based orthology analysis tool, SwiftOrtho, which is optimized for speed and memory usage when applied to large-scale data. SwiftOrtho uses long k-mers to speed up homology search, while using a reduced amino acid alphabet and spaced seeds to compensate for the loss of sensitivity due to long k-mers. In addition, it uses an affinity propagation algorithm to reduce the memory usage when clustering large-scale orthology relationships into orthologous groups. In our tests, SwiftOrtho was the only tool that completed orthology analysis of proteins from 1,760 bacterial genomes on a computer with only 4 GB RAM. Using various standard orthology data sets, we also show that SwiftOrtho has a high accuracy.ConclusionsSwiftOrtho enables the accurate comparative genomic analyses of thousands of genomes using low-memory computers. SwiftOrtho is available at https://github.com/Rinoahu/SwiftOrtho
- Research Article
12
- 10.1111/mpp.13043
- Mar 3, 2021
- Molecular Plant Pathology
Pathogens deploy a wide range of pathogenicity factors, including a plethora of proteases, to modify host tissue or manipulate host defences. Metalloproteases (MPs) have been implicated in virulence in several animal and plant pathogens. Here we investigated the repertoire of MPs in 46 stramenopile species including 37 oomycetes, 5 diatoms, and 4 brown algae. Screening their complete proteomes using hidden Markov models (HMMs) trained for MP detection resulted in over 4,000 MPs, with most species having between 65 and 100 putative MPs. Classification in clans and families according to the MEROPS database showed a highly diverse MP repertoire in each species. Analyses of domain composition, orthologous groups, distribution, and abundance within the stramenopile lineage revealed a few oomycete‐specific MPs and MPs potentially related to lifestyle. In‐depth analyses of MPs in the plant pathogen Phytophthora infestans revealed 91 MPs, divided over 21 protein families, including 25 MPs with a predicted signal peptide or signal anchor. Expression profiling showed different patterns of MP gene expression during pre‐infection and infection stages. When expressed in leaves of Nicotiana benthamiana, 12 MPs changed the sizes of lesions caused by inoculation with P. infestans; with 9 MPs the lesions were larger, suggesting a positive effect on the virulence of P. infestans, while 3 MPs had a negative effect, resulting in smaller lesions. To the best of our knowledge, this is the first systematic inventory of MPs in oomycetes and the first study pinpointing MPs as potential pathogenicity factors in Phytophthora.
- Research Article
62
- 10.3389/fmicb.2017.02389
- Dec 5, 2017
- Frontiers in Microbiology
Mycobacterium bovis causes bovine tuberculosis and is the main organism responsible for zoonotic tuberculosis in humans. We performed the sequencing, assembly and annotation of a Brazilian strain of M. bovis named SP38, and performed comparative genomics of M. bovis genomes deposited in GenBank. M. bovis SP38 has a traditional tuberculous mycobacterium genome of 4,347,648 bp, with 65.5% GC, and 4,216 genes. The majority of CDSs (2,805, 69.3%) have predictive function, while 1,206 (30.07%) are hypothetical. For comparative analysis, 31 M. bovis, 32 M. bovis BCG, and 23 Mycobacterium tuberculosis genomes available in GenBank were selected. M. bovis RDs (regions of difference) and Clonal Complexes (CC) were identified in silico. Genome dynamics of bacterial groups were analyzed by gene orthology and polymorphic sites identification. M. bovis polymorphic sites were used to construct a phylogenetic tree. Our RD analyses resulted in the exclusion of three genomes, mistakenly annotated as virulent M. bovis. M. bovis SP38 along with strain 35 represent the first report of CC European 2 in Brazil, whereas two other M. bovis strains failed to be classified within current CC. Results of M. bovis orthologous genes analysis suggest a process of genome remodeling through genomic decay and gene duplication. Quantification, pairwise comparisons and distribution analyses of polymorphic sites demonstrate greater genetic variability of M. tuberculosis when compared to M. bovis and M. bovis BCG (p ≤ 0.05), indicating that currently defined M. tuberculosis lineages are more genetically diverse than M. bovis CC and animal-adapted MTC (M. tuberculosis Complex) species. As expected, polymorphic sites annotation shows that M. bovis BCG are subjected to different evolutionary pressures when compared to virulent mycobacteria. Lastly, M. bovis phylogeny indicates that polymorphic sites may be used as markers of M. bovis lineages in association with CC. Our findings highlight the need to better understand host-pathogen co-evolution in genetically homogeneous and/or diverse host populations, considering the fact that M. bovis has a broader host range when compared to M. tuberculosis. Also, the identification of M. bovis genomes not classified within CC indicates that the diversity of M. bovis lineages may be larger than previously thought or that current classification should be reviewed.
- Research Article
3
- 10.3390/biom13071116
- Jul 13, 2023
- Biomolecules
Tandem repeats in proteins are patterns of residues repeated directly adjacent to each other. The evolution of these repeats can be assessed by using groups of homologous sequences, which can help pointing to events of unit duplication or deletion. High pressure in a protein family for variation of a given type of repeat might point to their function. Here, we propose the analysis of protein families to calculate protein short tandem repeats (pSTRs) in each protein sequence and assess their variability within the family in terms of number of units. To facilitate this analysis, we developed the pSTR tool, a method to analyze the evolution of protein short tandem repeats in a given protein family by pairwise comparisons between evolutionarily related protein sequences. We evaluated pSTR unit number variation in protein families of 12 complete metazoan proteomes. We hypothesize that families with more dynamic ensembles of repeats could reflect particular roles of these repeats in processes that require more adaptability.
- Research Article
2
- 10.1093/jhered/esaf077
- Oct 13, 2025
- The Journal of heredity
Bull kelp, Nereocystis luetkeana, is a northeastern Pacific kelp with a broad distribution from Alaska to central California. Its population declines have caused severe concerns in northern California, the Salish Sea in Washington, and recently in some populations in Oregon. Despite bull kelp's accumulated ecological and physiological studies, an assembled and annotated genomic reference was still unavailable. Here, we report the complete and annotated genome of N. luetkeana, produced by the California Conservation Genomics Project (CCGP), which aims to reveal genomic diversity patterns across California by sequencing the complete genomes of approximately 150 carefully selected species. The genome was assembled into 1,562 scaffolds with 449.82Mb, 80× of coverage, and 22,952 gene models. BUSCO assembly showed a completeness score of 72% for the stramenopiles gene set. The mitochondria and chloroplast genome sequences have 37 Kb and 131Mb, respectively. The orthology analysis between 10 Phaeophycean genomes showed 1,065 expanded and 286 unique orthogroups for this species. Pairwise comparisons showed 542 orthogroups present only in N. luetkeana and Macrocystis pyrifera, another large-body kelp. The enrichment analysis of these orthogroups showed important functions related to central metabolism and signaling due to ATPase enrichment in these two species. This genome assembly will provide an essential resource for the ecology, evolution, conservation, and breeding of bull kelp.
- Research Article
- 10.5498/wjp.v16.i2.111012
- Feb 19, 2026
- World Journal of Psychiatry
BACKGROUNDAutism spectrum disorder (ASD) is a neurodevelopmental disorder characterized by pronounced behavioral heterogeneity and individual variability. Growing evidence indicates a strong association between gut microbiota and ASD; however, differences in microbial functions across varying levels of ASD severity remain poorly understood. Monozygotic twins (MZs) provide an appropriate model for examining the influence of nonshared environmental factors in ASD.AIMTo investigate the effects of the gut microbiome in MZs with ASD using 16S ribosomal RNA sequencing.METHODSParticipants were recruited from the Chinese MZs with autism spectrum disorder (MZCo-ASD) cohort and stratified into mild MZCo-ASD and severe MZCo-ASD (MZCo-ASD-H) groups based on their Childhood Autism Rating Scale scores. Fecal samples were collected and analyzed using 16S ribosomal RNA sequencing.RESULTSAlthough overall microbial diversity did not differ significantly between the groups, gut microbiota composition was notably altered. At the genus level, Porphyromonas was significantly enriched in the MZCo-ASD-H group. Clusters of Orthologous Groups analysis revealed decreased expression of key genes in the MZCo-ASD-H group, including fructose-1,6-bisphosphatase, membrane-bound lytic murein transglycosylase, PasI (part of the RatAB toxin-antitoxin system), HmoA, and a glycoside hydrolase family 25 domain-containing protein. Kyoto Encyclopedia of Genes and Genomes Orthology analysis showed that msmF (K10118) and msmG (K10119), involved in oligosaccharide transport, were significantly downregulated in the MZCo-ASD-H group, suggesting a reduced microbial capacity for prebiotic carbohydrate utilization.CONCLUSIONDespite similar overall diversity, children with severe ASD exhibited distinct gut microbiota structures and functional impairments. The enrichment of Porphyromonas, along with the reduced expression of genes involved in carbohydrate metabolism and stress responses in the high-severity group, suggests an association between gut microbial dysregulation and ASD severity. These findings provide new insights into microbiota-related mechanisms underlying ASD and highlight potential functional targets for intervention.
- Research Article
112
- 10.1186/1471-2164-14-552
- Jan 1, 2013
- BMC Genomics
BackgroundLitchi (Litchi chinensis Sonn.) is one of the most important fruit trees cultivated in tropical and subtropical areas. However, a lack of transcriptomic and genomic information hinders our understanding of the molecular mechanisms underlying fruit set and fruit development in litchi. Shading during early fruit development decreases fruit growth and induces fruit abscission. Here, high-throughput RNA sequencing (RNA-Seq) was employed for the de novo assembly and characterization of the fruit transcriptome in litchi, and differentially regulated genes, which are responsive to shading, were also investigated using digital transcript abundance(DTA)profiling.ResultsMore than 53 million paired-end reads were generated and assembled into 57,050 unigenes with an average length of 601 bp. These unigenes were annotated by querying against various public databases, with 34,029 unigenes found to be homologous to genes in the NCBI GenBank database and 22,945 unigenes annotated based on known proteins in the Swiss-Prot database. In further orthologous analyses, 5,885 unigenes were assigned with one or more Gene Ontology terms, 10,234 hits were aligned to the 24 Clusters of Orthologous Groups classifications and 15,330 unigenes were classified into 266 Kyoto Encyclopedia of Genes and Genomes pathways. Based on the newly assembled transcriptome, the DTA profiling approach was applied to investigate the differentially expressed genes related to shading stress. A total of 3.6 million and 3.5 million high-quality tags were generated from shaded and non-shaded libraries, respectively. As many as 1,039 unigenes were shown to be significantly differentially regulated. Eleven of the 14 differentially regulated unigenes, which were randomly selected for more detailed expression comparison during the course of shading treatment, were identified as being likely to be involved in the process of fruitlet abscission in litchi.ConclusionsThe assembled transcriptome of litchi fruit provides a global description of expressed genes in litchi fruit development, and could serve as an ideal repository for future functional characterization of specific genes. The DTA analysis revealed that more than 1000 differentially regulated unigenes respond to the shading signal, some of which might be involved in the fruitlet abscission process in litchi, shedding new light on the molecular mechanisms underlying organ abscission.
- Research Article
10
- 10.1089/cmb.2015.0084
- Jul 10, 2015
- Journal of Computational Biology
Identification and clustering of orthologous genes plays an important role in developing evolutionary models such as validating convergent and divergent phylogeny and predicting functional proteins in newly sequenced species of unverified nucleotide protein mappings. Here, we introduce an application of subspace clustering as applied to orthologous gene sequences and discuss the initial results. The working hypothesis is based upon the concept that genetic changes between nucleotide sequences coding for proteins among selected species and groups may lie within a union of subspaces for clusters of the orthologous groups. Estimates for the subspace dimensions were computed for a small population sample. A series of experiments was performed to cluster randomly selected sequences. The experimental design allows for both false positives and false negatives, and estimates for the statistical significance are provided. The clustering results are consistent with the main hypothesis. A simple random mutation binary tree model is used to simulate speciation events that show the interdependence of the subspace rank versus time and mutation rates. The simple mutation model is found to be largely consistent with the observed subspace clustering singular value results. Our study indicates that the subspace clustering method may be applied in orthology analysis.
- Research Article
75
- 10.3389/fgene.2017.00165
- Oct 31, 2017
- Frontiers in Genetics
Nowadays defying homology relationships among sequences is essential for biological research. Within homology the analysis of orthologs sequences is of great importance for computational biology, annotation of genomes and for phylogenetic inference. Since 2007, with the increase in the number of new sequences being deposited in large biological databases, researchers have begun to analyse computerized methodologies and tools aimed at selecting the most promising ones in the prediction of orthologous groups. Literature in this field of research describes the problems that the majority of available tools show, such as those encountered in accuracy, time required for analysis (especially in light of the increasing volume of data being submitted, which require faster techniques) and the automatization of the process without requiring manual intervention. Conducting our search through BMC, Google Scholar, NCBI PubMed, and Expasy, we examined more than 600 articles pursuing the most recent techniques and tools developed to solve most the problems still existing in orthology detection. We listed the main computational tools created and developed between 2011 and 2017, taking into consideration the differences in the type of orthology analysis, outlining the main features of each tool and pointing to the problems that each one tries to address. We also observed that several tools still use as their main algorithm the BLAST “all-against-all” methodology, which entails some limitations, such as limited number of queries, computational cost, and high processing time to complete the analysis. However, new promising tools are being developed, like OrthoVenn (which uses the Venn diagram to show the relationship of ortholog groups generated by its algorithm); or proteinOrtho (which improves the accuracy of ortholog groups); or ReMark (tackling the integration of the pipeline to turn the entry process automatic); or OrthAgogue (using algorithms developed to minimize processing time); and proteinOrtho (developed for dealing with large amounts of biological data). We made a comparison among the main features of four tool and tested them using four for prokaryotic genomas. We hope that our review can be useful for researchers and will help them in selecting the most appropriate tool for their work in the field of orthology.
- Research Article
20
- 10.1016/j.fm.2019.103392
- Nov 26, 2019
- Food Microbiology
Investigation of genomic characteristics and carbohydrates’ metabolic activity of Lactococcus lactis subsp. lactis during ripening of a Swiss-type cheese
- Research Article
- 10.3389/fmicb.2025.1494490
- May 1, 2025
- Frontiers in microbiology
The global rise in antibiotic resistance and emergence of new bacterial pathogens pose a significant threat to public health. Novel approaches to uncover potential novel diagnostic and therapeutic targets for these pathogens are needed. In this study, we conducted a large-scale, phylogenetic-based orthology analysis (OA) to compare the proteomes of pathogenic to humans (HP) and non-pathogenic to humans (NHP) bacterial strains across 734 strains from 514 species and 91 families. Using a dedicated workflow, we identified 4,383 hierarchical orthologous groups (HOGs) significantly associated with the HP label, many of which are linked to critical factors such as stress tolerance, metabolic versatility, and antibiotic resistance. Both known virulence factors (VFs) and potential novel widespread pathogenicity determinants were uncovered, supported by both statistical testing and complementary protein domain analysis. By integrating curated strain-level pathogenicity annotations from BacSPaD with phylogeny-based OA, we introduce a novel approach and provide a novel resource for bacterial pathogenicity research.
- Research Article
142
- 10.1101/gr.209402
- Apr 1, 2002
- Genome Research
All protein sequences from 19 complete chloroplast genomes (cpDNA) have been studied using a new computational method able to analyze functional correlations among series of protein sequences contained in complete proteomes. First, all open reading frames (ORFs) from the cpDNAs, comprising a total of 2266 protein sequences, were compared against the 3168 proteins from Synechocystis PCC6803 complete genome to find functionally related orthologous proteins. Additionally, all cpDNA genomes were pairwise compared to find orthologous groups not present in cyanobacteria. Annotations in the cluster of othologous proteins database and CyanoBase were used as reference for the functional assignments. Following this protocol, new functional assignments were made for ORFs of unknown function and for ycfs (hypothetical chloroplast frames), which still lack a functional assignment. Using this information, a matrix of functional relationships was derived from profiles of the presence and/or absence of orthologous proteins; the matrix included 1837 proteins in 277 orthologous clusters. A factor analysis study of this matrix, followed by cluster analysis, allowed us to obtain accurate phylogenetic reconstructions and the detection of genes probably involved in speciation as phylogenetic correlates. Finally, by grouping common evolutionary patterns, we show that it is possible to determine functionally linked protein networks. This has allowed us to suggest putative associations for some unknown ORFs.