Accelerate Literature Icon
Want to do a literature review? Try our new Literature Review workflow

Comparative performance of structural aligners in functional domain annotation.

  • Abstract
  • Literature Map
  • Similar Papers
Abstract
Translate article icon Translate Article Star icon

Accurate protein domain annotation is essential for inferring protein function, and databases such as Pfam provide sequence-derived signatures for thousands of domain families. Because protein structure is more evolutionarily conserved than sequence, structure-based searches can detect homologous relationships even at low sequence identity (typically below 30%), where pairwise sequence aligners often lose sensitivity. Here, we leverage AlphaFold-derived structures of Pfam domain instances to systematically evaluate structure-based versus sequence-based methods for Pfam annotation. We benchmarked three structural aligners (Reseek, Foldseek, TM-align) against sequence-based methods (MMseqs, HMMER) using both exhaustive all-against-all searches and a split-family design that enables direct comparison of pairwise and profile-based ranking performance. We also evaluated residue-level alignment accuracy using Pfam multiple sequence alignments as reference and investigated whether profile-derived information can improve structural hit ranking. In all-against-all searches, Reseek achieved the highest sensitivity up to the first false positive (AUC = 0.85), outperforming Foldseek (0.81), TM-align (0.76), and MMseqs (0.46). In split-family evaluation, HMMER remained superior (maximum F1=0.991), highlighting the continued strength of sequence-profile approaches for family-level annotation. Performance varied substantially across domain families, with average sequence identity emerging as the strongest predictor of success. Structural aligners consistently produced more accurate residue-level mappings than pairwise sequence methods. Finally, incorporating profile-derived information via rescoring improved structural annotation performance for short domains, suggesting a path toward profile-informed structure-based domain annotation.

Similar Papers
  • PDF Download Icon
  • Research Article
  • Cite Count Icon 195
  • 10.1074/jbc.m414508200
Crystal Structure of Vinorine Synthase, the First Representative of the BAHD Superfamily
  • Apr 1, 2005
  • Journal of Biological Chemistry
  • Xueyan Ma + 4 more

Vinorine synthase is an acetyltransferase that occupies a central role in the biosynthesis of the antiarrhythmic monoterpenoid indole alkaloid ajmaline in the plant Rauvolfia. Vinorine synthase belongs to the benzylalcohol acetyl-, anthocyanin-O-hydroxy-cinnamoyl-, anthranilate-N-hydroxy-cinnamoyl/benzoyl-, deacetylvindoline acetyltransferase (BAHD) enzyme superfamily, members of which are involved in the biosynthesis of several important drugs, such as morphine, Taxol, or vindoline, a precursor of the anti-cancer drugs vincaleucoblastine and vincristine. The x-ray structure of vinorine synthase is described at 2.6-angstrom resolution. Despite low sequence identity, the two-domain structure of vinorine synthase shows surprising similarity with structures of several CoA-dependent acyltransferases such as dihydrolipoyl transacetylase, polyketide-associated protein A5, and carnitine acetyltransferase. All conserved residues typical for the BAHD family are found in domain 1. His160 of the HXXXD motif functions as a general base during catalysis. It is located in the center of the reaction channel at the interface of both domains and is accessible from both sides. The channel runs through the entire molecule, allowing the substrate and co-substrate to bind independently. Asp164 points away from the catalytic site and seems to be of structural rather than catalytic importance. Surprisingly, the DFGWG motif, which is indispensable for the catalyzed reaction and unique to the BAHD family, is located far away from the active site and seems to play only a structural role. Vinorine synthase represents the first solved protein structure of the BAHD superfamily.

  • Research Article
  • 10.1109/bibm.2011.74
R-PASS: A Fast Structure-based RNA Sequence Alignment Algorithm.
  • Nov 1, 2011
  • Proceedings. IEEE International Conference on Bioinformatics and Biomedicine
  • Yanan Jiang + 4 more

We present a fast pairwise RNA sequence alignment method using structural information, named R-PASS (RNA Pairwise Alignment of Structure and Sequence), which shows good accuracy on sequences with low sequence identity and significantly faster than alternative methods. The method begins by representing RNA secondary structure as a set of structure motifs. The motifs from two RNAs are then used as input into a bipartite graph-matching algorithm, which determines the structure matches. The matches are then used as constraints in a constrained dynamic programming sequence alignment procedure. The R-PASS method has an O(nm) complexity. We compare our method with two other structure-based alignment methods, LARA and ExpaLoc, and with a sequence-based alignment method, MAFFT, across three benchmarks and obtain favorable results in accuracy and orders of magnitude faster in speed.

  • Research Article
  • Cite Count Icon 17
  • 10.1016/j.celrep.2022.111030
An anti-picornaviral strategy based on the crystal structure of foot-and-mouth disease virus 2C protein.
  • Jul 1, 2022
  • Cell Reports
  • Chu Zhang + 10 more

An anti-picornaviral strategy based on the crystal structure of foot-and-mouth disease virus 2C protein.

  • Research Article
  • Cite Count Icon 261
  • 10.1002/(sici)1097-0134(20000701)40:1<6::aid-prot30>3.0.co;2-7
Large-scale comparison of protein sequence alignment algorithms with structure alignments.
  • May 11, 2000
  • Proteins: Structure, Function, and Genetics
  • J Michael Sauder + 2 more

Sequence alignment programs such as BLAST and PSI-BLAST are used routinely in pairwise, profile-based, or intermediate-sequence-search (ISS) methods to detect remote homologies for the purposes of fold assignment and comparative modeling. Yet, the sequence alignment quality of these methods at low sequence identity is not known. We have used the CE structure alignment program (Shindyalov and Bourne, Prot Eng 1998;11:739) to derive sequence alignments for all superfamily and family-level related proteins in the SCOP domain database. CE aligns structures and their sequences based on distances within each protein, rather than on interprotein distances. We compared BLAST, PSI-BLAST, CLUSTALW, and ISS alignments with the CE structural alignments. We found that global alignments with CLUSTALW were very poor at low sequence identity (<25%), as judged by the CE alignments. We used PSI-BLAST to search the nonredundant sequence database (nr) with every sequence in SCOP using up to four iterations. The resulting matrix was used to search a database of SCOP sequences. PSI-BLAST is only slightly better than BLAST in alignment accuracy on a per-residue basis, but PSI-BLAST matrix alignments are much longer than BLAST's, and so align correctly a larger fraction of the total number of aligned residues in the structure alignments. Any two SCOP sequences in the same superfamily that shared a hit or hits in the nr PSI-BLAST searches were identified as linked by the shared intermediate sequence. We examined the quality of the longest SCOP-query/ SCOP-hit alignment via an intermediate sequence, and found that ISS produced longer alignments than PSI-BLAST searches alone, of nearly comparable per-residue quality. At 10-15% sequence identity, BLAST correctly aligns 28%, PSI-BLAST 40%, and ISS 46% of residues according to the structure alignments. We also compared CE structure alignments with FSSP structure alignments generated by the DALI program. In contrast to the sequence methods, CE and structure alignments from the FSSP database identically align 75% of residue pairs at the 10-15% level of sequence identity, indicating that there is substantial room for improvement in these sequence alignment methods. BLAST produced alignments for 8% of the 10,665 nonimmunoglobulin SCOP superfamily sequence pairs (nearly all <25% sequence identity), PSI-BLAST matched 17% and the double-PSI-BLAST ISS method aligned 38% with E-values <10.0. The results indicate that intermediate sequences may be useful not only in fold assignment but also in achieving more complete sequence alignments for comparative modeling.

  • Research Article
  • Cite Count Icon 173
  • 10.1110/ps.051892906
A comparative study of available software for high-accuracy homology modeling: from sequence alignments to structural models.
  • Apr 1, 2006
  • Protein Science
  • Akbar Nayeem + 2 more

An open question in protein homology modeling is, how well do current modeling packages satisfy the dual criteria of quality of results and practical ease of use? To address this question objectively, we examined homology-built models of a variety of therapeutically relevant proteins. The sequence identities across these proteins range from 19% to 76%. A novel metric, the difference alignment index (DAI), is developed to aid in quantifying the quality of local sequence alignments. The DAI is also used to construct the relative sequence alignment (RSA), a new representation of global sequence alignment that facilitates comparison of sequence alignments from different methods. Comparisons of the sequence alignments in terms of the RSA and alignment methodologies are made to better understand the advantages and caveats of each method. All sequence alignments and corresponding 3D models are compared to their respective structure-based alignments and crystal structures. A variety of protein modeling software was used. We find that at sequence identities >40%, all packages give similar (and satisfactory) results; at lower sequence identities (<25%), the sequence alignments generated by Profit and Prime, which incorporate structural information in their sequence alignment, stand out from the rest. Moreover, the model generated by Prime in this low sequence identity region is noted to be superior to the rest. Additionally, we note that DSModeler and MOE, which generate reasonable models for sequence identities >25%, are significantly more functional and easier to use when compared with the other structure-building software.

  • Research Article
  • Cite Count Icon 6
  • 10.1002/prot.24134
Overcoming sequence misalignments with weighted structural superposition
  • Jul 28, 2012
  • Proteins: Structure, Function, and Bioinformatics
  • Nickolay A Khazanov + 3 more

An appropriate structural superposition identifies similarities and differences between homologous proteins that are not evident from sequence alignments alone. We have coupled our Gaussian-weighted RMSD (wRMSD) tool with a sequence aligner and seed extension (SE) algorithm to create a robust technique for overlaying structures and aligning sequences of homologous proteins (HwRMSD). HwRMSD overcomes errors in the initial sequence alignment that would normally propagate into a standard RMSD overlay. SE can generate a corrected sequence alignment from the improved structural superposition obtained by wRMSD. HwRMSD's robust performance and its superiority over standard RMSD are demonstrated over a range of homologous proteins. Its better overlay results in corrected sequence alignments with good agreement to HOMSTRAD. Finally, HwRMSD is compared to established structural alignment methods: FATCAT, secondary-structure matching, combinatorial extension, and Dalilite. Most methods are comparable at placing residue pairs within 2 Å, but HwRMSD places many more residue pairs within 1 Å, providing a clear advantage. Such high accuracy is essential in drug design, where small distances can have a large impact on computational predictions. This level of accuracy is also needed to correct sequence alignments in an automated fashion, especially for omics-scale analysis. HwRMSD can align homologs with low-sequence identity and large conformational differences, cases where both sequence-based and structural-based methods may fail. The HwRMSD pipeline overcomes the dependency of structural overlays on initial sequence pairing and removes the need to determine the best sequence-alignment method, substitution matrix, and gap parameters for each unique pair of homologs.

  • Research Article
  • Cite Count Icon 12
  • 10.1093/molbev/msaf149
Newly Developed Structure-Based Methods Do Not Outperform Standard Sequence-Based Methods for Large-Scale Phylogenomics.
  • Jun 23, 2025
  • Molecular biology and evolution
  • Giacomo Mutti + 2 more

Recent developments in protein structure prediction have allowed the use of this previously limited source of information at genome-wide scales. It has been proposed that the use of structural information may offer advantages over sequences in phylogenetic reconstruction, due to their slower rate of evolution and direct correlation to function. Here, we examined how recently developed methods for structure-based homology search and tree reconstruction compare with current state-of-the-art sequence-based methods in reconstructing genome-wide collections of gene phylogenies (i.e. phylomes). While structure-based methods can be useful in specific scenarios, we found that their current performance does not justify using the newly developed structure-based methods as a default choice in large-scale phylogenetic studies. On the one hand, the best performing sequence-based tree reconstruction methods still outperform structure-based methods for this task. On the other hand, structure-based homology detection methods provide larger lists of candidate homologs, as previously reported. However, this comes at the expense of missing hits identified by sequence-based methods, as well as providing sets of homolog candidates with higher fractions of false positives. These insights help to guide the use of structural data in comparative genomics and highlight the need to continue improving structure-based approaches. Our pipeline is fully reproducible and has been implemented in a Snakemake workflow. This will facilitate a continuous assessment of future improvements of structure-based tools in the AlphaFold era.

  • Research Article
  • Cite Count Icon 34
  • 10.1007/s10969-012-9126-6
Structure- and sequence-based function prediction for non-homologous proteins
  • Jan 22, 2012
  • Journal of Structural and Functional Genomics
  • Lee Sael + 2 more

The structural genomics projects have been accumulating an increasing number of protein structures, many of which remain functionally unknown. In parallel effort to experimental methods, computational methods are expected to make a significant contribution for functional elucidation of such proteins. However, conventional computational methods that transfer functions from homologous proteins do not help much for these uncharacterized protein structures because they do not have apparent structural or sequence similarity with the known proteins. Here, we briefly review two avenues of computational function prediction methods, i.e. structure-based methods and sequence-based methods. The focus is on our recent developments of local structure-based and sequence-based methods, which can effectively extract function information from distantly related proteins. Two structure-based methods, Pocket-Surfer and Patch-Surfer, identify similar known ligand binding sites for pocket regions in a query protein without using global protein fold similarity information. Two sequence-based methods, protein function prediction and extended similarity group, make use of weakly similar sequences that are conventionally discarded in homology based function annotation. Combined together with experimental methods we hope that computational methods will make leading contribution in functional elucidation of the protein structures.

  • Research Article
  • Cite Count Icon 7
  • 10.1371/journal.pcbi.1012526
Structure-aware annotation of leucine-rich repeat domains.
  • Nov 5, 2024
  • PLoS computational biology
  • Boyan Xu + 4 more

Protein domain annotation is typically done by predictive models such as HMMs trained on sequence motifs. However, sequence-based annotation methods are prone to error, particularly in calling domain boundaries and motifs within them. These methods are limited by a lack of structural information accessible to the model. With the advent of deep learning-based protein structure prediction, existing sequenced-based domain annotation methods can be improved by taking into account the geometry of protein structures. We develop dimensionality reduction methods to annotate repeat units of the Leucine Rich Repeat solenoid domain. The methods are able to correct mistakes made by existing machine learning-based annotation tools and enable the automated detection of hairpin loops and structural anomalies in the solenoid. The methods are applied to 127 predicted structures of LRR-containing intracellular innate immune proteins in the model plant Arabidopsis thaliana and validated against a benchmark dataset of 172 manually-annotated LRR domains.

  • Research Article
  • 10.1101/2023.10.27.562987
Structure-Aware Annotation of Leucine-rich Repeat Domains
  • Nov 1, 2023
  • bioRxiv
  • Boyan Xu + 4 more

Protein domain annotation is typically done by predictive models such as HMMs trained on sequence motifs. However, sequence-based annotation methods are prone to error, particularly in calling domain boundaries and motifs within them. These methods are limited by a lack of structural information accessible to the model. With the advent of deep learning-based protein structure prediction, existing sequenced-based domain annotation methods can be improved by taking into account the geometry of protein structures. We develop dimensionality reduction methods to annotate repeat units of the Leucine Rich Repeat solenoid domain. The methods are able to correct mistakes made by existing machine learning-based annotation tools and enable the automated detection of hairpin loops and structural anomalies in the solenoid. The methods are applied to 127 predicted structures of LRR-containing intracellular innate immune proteins in the model plant Arabidopsis thaliana and validated against a benchmark dataset of 172 manually-annotated LRR domains.

  • Research Article
  • Cite Count Icon 587
  • 10.1006/jmbi.1998.2221
Sequence comparisons using multiple sequences detect three times as many remote homologues as pairwise methods
  • Dec 1, 1998
  • Journal of Molecular Biology
  • Jong Park + 6 more

Sequence comparisons using multiple sequences detect three times as many remote homologues as pairwise methods

  • Research Article
  • Cite Count Icon 10
  • 10.1007/bf02407306
Phyletic relationships of protein structures based on spatial preference of residues
  • Jan 1, 1993
  • Journal of Molecular Evolution
  • Chun-Xu Qu + 3 more

A structure-based scoring matrix MDPRE was derived from amino acid spatial preferences in protein structures. Sequence alignment and evolutionary studies by using MDPRE matrix gave similar results as those from ordinary sequence and structure alignments. It is interesting that a matrix derived from structure data solely could give comparable alignment results, strongly indicating the intimate connection between protein sequences and structures. The branch order and length from this approach were close to those obtained by a structure comparison method. Thus, by applying this structure-based matrix, the trees obtained should reflect evolutionary characteristics of protein structure. This approach takes advantage over a direct structure comparison in that (1) only a sequence and MDPRE matrix are needed, making it simple and widely applicable (especially in the absence of 3-dimensional protein structure data); (2) an established algorithm for sequence alignment and tree building could be employed, providing opportunities for direct comparison between matrices from different methodologies. One of the most striking features of this method is its capability to detect protein structure homologies when the sequence identities are low. This was well reflected in the given examples of the alignment of dinucleotide-binding domains.

  • Research Article
  • Cite Count Icon 4
  • 10.2165/00822942-200403020-00009
MSAT
  • Jan 1, 2004
  • Applied Bioinformatics
  • Te Ren + 3 more

This article describes the development of a new method for multiple sequence alignment based on fold-level protein structure alignments, which provides an improvement in accuracy compared with the most commonly used sequence-only-based techniques. This method integrates the widely used, progressive multiple sequence alignment approach ClustalW with the Topology of Protein Structure (TOPS) topology-based alignment algorithm. The TOPS approach produces a structural alignment for the input protein set by using a topology-based pattern discovery program, providing a set of matched sequence regions that can be used to guide a sequence alignment using ClustalW. The resulting alignments are more reliable than a sequence-only alignment, as determined by 20-fold cross-validation with a set of 106 protein examples from the CATH database, distributed in seven superfold families. The method is particularly effective for sets of proteins that have similar structures at the fold level but low sequence identity. The aim of this research is to contribute towards bridging the gap between protein sequence and structure analysis, in the hope that this can be used to assist the understanding of the relationship between sequence, structure and function. The tool is available at http://balabio.dcs.gla.ac.uk/msat/.

  • Research Article
  • Cite Count Icon 34
  • 10.1002/prot.20843
Nontoxic crystal protein from Bacillus thuringiensis demonstrates a remarkable structural similarity to β‐pore‐forming toxins
  • Jan 6, 2006
  • Proteins: Structure, Function, and Bioinformatics
  • Toshihiko Akiba + 7 more

Nontoxic crystal protein from <i>Bacillus thuringiensis</i> demonstrates a remarkable structural similarity to β‐pore‐forming toxins

  • Research Article
  • Cite Count Icon 32
  • 10.1093/bioinformatics/btad646
THPLM: a sequence-based deep learning framework for protein stability changes prediction upon point variations using pretrained protein language model.
  • Oct 24, 2023
  • Bioinformatics (Oxford, England)
  • Jianting Gong + 10 more

Quantitative determination of protein thermodynamic stability is a critical step in protein and drug design. Reliable prediction of protein stability changes caused by point variations contributes to developing-related fields. Over the past decades, dozens of structure-based and sequence-based methods have been proposed, showing good prediction performance. Despite the impressive progress, it is necessary to explore wild-type and variant protein representations to address the problem of how to represent the protein stability change in view of global sequence. With the development of structure prediction using learning-based methods, protein language models (PLMs) have shown accurate and high-quality predictions of protein structure. Because PLM captures the atomic-level structural information, it can help to understand how single-point variations cause functional changes. Here, we proposed THPLM, a sequence-based deep learning model for stability change prediction using Meta's ESM-2. With ESM-2 and a simple convolutional neural network, THPLM achieved comparable or even better performance than most methods, including sequence-based and structure-based methods. Furthermore, the experimental results indicate that the PLM's ability to generate representations of sequence can effectively improve the ability of protein function prediction. The source code of THPLM and the testing data can be accessible through the following links: https://github.com/FPPGroup/THPLM.

Save Icon
Up Arrow
Open/Close
Notes

Save Important notes in documents

Highlight text to save as a note, or write notes directly

You can also access these Documents in Paperpal, our AI writing tool

Powered by our AI Writing Assistant