Protein Structure Similarity Research Articles

BackgroundAccurate sequence alignments are essential for homology searches and for building three-dimensional structural models of proteins. Since structure is better conserved than sequence, structure alignments have been used to guide sequence alignments and are commonly used as the gold standard for sequence alignment evaluation. Nonetheless, as far as we know, there is no report of a systematic evaluation of pairwise structure alignment programs in terms of the sequence alignment accuracy.ResultsIn this study, we evaluate CE, DaliLite, FAST, LOCK2, MATRAS, SHEBA and VAST in terms of the accuracy of the sequence alignments they produce, using sequence alignments from NCBI's human-curated Conserved Domain Database (CDD) as the standard of truth. We find that 4 to 9% of the residues on average are either not aligned or aligned with more than 8 residues of shift error and that an additional 6 to 14% of residues on average are misaligned by 1–8 residues, depending on the program and the data set used. The fraction of correctly aligned residues generally decreases as the sequence similarity decreases or as the RMSD between the Cα positions of the two structures increases. It varies significantly across CDD superfamilies whether shift error is allowed or not. Also, alignments with different shift errors occur between proteins within the same CDD superfamily, leading to inconsistent alignments between superfamily members. In general, residue pairs that are more than 3.0 Å apart in the reference alignment are heavily (>= 25% on average) misaligned in the test alignments. In addition, each method shows a different pattern of relative weaknesses for different SCOP classes. CE gives relatively poor results for β-sheet-containing structures (all-β, α/β, and α+β classes), DaliLite for "others" class where all but the major four classes are combined, and LOCK2 and VAST for all-β and "others" classes.ConclusionWhen the sequence similarity is low, structure-based methods produce better sequence alignments than by using sequence similarities alone. However, current structure-based methods still mis-align 11–19% of the conserved core residues when compared to the human-curated CDD alignments. The alignment quality of each program depends on the protein structural type and similarity, with DaliLite showing the most agreement with CDD on average.

Read full abstract

BackgroundOwing to rapid expansion of protein structure databases in recent years, methods of structure comparison are becoming increasingly effective and important in revealing novel information on functional properties of proteins and their roles in the grand scheme of evolutionary biology. Currently, the structural similarity between two proteins is measured by the root-mean-square-deviation (RMSD) in their best-superimposed atomic coordinates. RMSD is the golden rule of measuring structural similarity when the structures are nearly identical; it, however, fails to detect the higher order topological similarities in proteins evolved into different shapes. We propose new algorithms for extracting geometrical invariants of proteins that can be effectively used to identify homologous protein structures or topologies in order to quantify both close and remote structural similarities.ResultsWe measure structural similarity between proteins by correlating the principle components of their secondary structure interaction matrix. In our approach, the Principle Component Correlation (PCC) analysis, a symmetric interaction matrix for a protein structure is constructed with relationship parameters between secondary elements that can take the form of distance, orientation, or other relevant structural invariants. When using a distance-based construction in the presence or absence of encoded N to C terminal sense, there are strong correlations between the principle components of interaction matrices of structurally or topologically similar proteins.ConclusionThe PCC method is extensively tested for protein structures that belong to the same topological class but are significantly different by RMSD measure. The PCC analysis can also differentiate proteins having similar shapes but different topological arrangements. Additionally, we demonstrate that when using two independently defined interaction matrices, comparison of their maximum eigenvalues can be highly effective in clustering structurally or topologically similar proteins. We believe that the PCC analysis of interaction matrix is highly flexible in adopting various structural parameters for protein structure comparison.

Read full abstract

Protein Structure Similarity Research Articles

Related Topics

Articles published on Protein Structure Similarity

TOPS++FATCAT: Fast flexible structural alignment using constraints derived from TOPS+ Strings Model

Protein structure alignment considering phenotypic plasticity

Statistics of Random Protein Superpositions: p-Values for Pairwise Structure Alignment

Synthesis of a dysidiolide-inspired compound library and discovery of acetylcholinesterase inhibitors based on protein structure similarity clustering (PSSC)

BIOS: Similarity-based design of natural product derived compound collections

Accuracy of structure-based sequence alignment of automatic methods

Multivariate statistical analysis of large-scale IgE antibody measurements reveals allergen extract relationships in sensitized individuals

Protein structural similarity search by Ramachandran codes

Reticulon-like proteins in Arabidopsis thaliana: Structural organization and ER localization

Knowledge-Driven Multidimensional Indexing Structure for Biomedical Media Database Retrieval

Chemometric Approach in Quantification of Structural Identity/Similarity of Proteins in Biopharmaceuticals

PROTMAP2D: visualization, comparison and analysis of 2D maps of protein structure

Detection of Legume Proteins Cross-reactivity by Immunoblot Using Human Plasma of Individuals with Food Allergies to Peanut and/or Soybean

Protein Structure Similarity Clustering: Dynamic Treatment of PDB Structures Facilitates Clustering

The Ig Doublet Z1Z2: A Model System for the Hybrid Analysis of Conformational Dynamics in Ig Tandems from Titin

StructSorter: A Method for Continuously Updating a Comprehensive Protein Structure Alignment Database

Charting Biological and Chemical Space: PSSC and SCONP as Guiding Principles for the Development of Compound Collections Based on Natural Product Scaffolds

Protein structure similarity from Principle Component Correlation analysis.

Flexible Structural Neighborhood--a database of protein structural similarities and alignments

Charting biologically relevant chemical space: A structural classification of natural products (SCONP)

Lead the way for us

Editage

Paperpal

R Discovery

Mind the Graph

Protein Structure Similarity Research Articles

Related Topics

Articles published on Protein Structure Similarity

TOPS++FATCAT: Fast flexible structural alignment using constraints derived from TOPS+ Strings Model

Protein structure alignment considering phenotypic plasticity

Statistics of Random Protein Superpositions: p-Values for Pairwise Structure Alignment

Synthesis of a dysidiolide-inspired compound library and discovery of acetylcholinesterase inhibitors based on protein structure similarity clustering (PSSC)

BIOS: Similarity-based design of natural product derived compound collections

Accuracy of structure-based sequence alignment of automatic methods

Multivariate statistical analysis of large-scale IgE antibody measurements reveals allergen extract relationships in sensitized individuals

Protein structural similarity search by Ramachandran codes

Reticulon-like proteins in Arabidopsis thaliana: Structural organization and ER localization

Knowledge-Driven Multidimensional Indexing Structure for Biomedical Media Database Retrieval

Chemometric Approach in Quantification of Structural Identity/Similarity of Proteins in Biopharmaceuticals

PROTMAP2D: visualization, comparison and analysis of 2D maps of protein structure

Detection of Legume Proteins Cross-reactivity by Immunoblot Using Human Plasma of Individuals with Food Allergies to Peanut and/or Soybean

Protein Structure Similarity Clustering: Dynamic Treatment of PDB Structures Facilitates Clustering

The Ig Doublet Z1Z2: A Model System for the Hybrid Analysis of Conformational Dynamics in Ig Tandems from Titin

StructSorter: A Method for Continuously Updating a Comprehensive Protein Structure Alignment Database

Charting Biological and Chemical Space: PSSC and SCONP as Guiding Principles for the Development of Compound Collections Based on Natural Product Scaffolds

Protein structure similarity from Principle Component Correlation analysis.

Flexible Structural Neighborhood--a database of protein structural similarities and alignments

Charting biologically relevant chemical space: A structural classification of natural products (SCONP)