Large scale analysis of protein‐binding cavities using self‐organizing maps and wavelet‐based surface patches to describe functional properties, selectivity discrimination, and putative cross‐reactivity

Katrin Kupas,Gerhard Klebe,Alfred Ultsch

doi:10.1002/prot.21823

Abstract

A new method to discover similar substructures in protein binding pockets, independently of sequence and folding patterns or secondary structure elements, is introduced. The solvent-accessible surface of a binding pocket, automatically detected as a depression on the protein surface, is divided into a set of surface patches. Each surface patch is characterized by its shape as well as by its physicochemical characteristics. Wavelets defined on surfaces are used for the description of the shape, as they have the great advantage of allowing a comparison at different resolutions. The number of coefficients to describe the wavelets can be chosen with respect to the size of the considered data set. The physicochemical characteristics of the patches are described by the assignment of the exposed amino acid residues to one or more of five different properties determinant for molecular recognition. A self-organizing neural network is used to project the high-dimensional feature vectors onto a two-dimensional layer of neurons, called a map. To find similarities between the binding pockets, in both geometrical and physicochemical features, a clustering of the projected feature vector is performed using an automatic distance- and density-based clustering algorithm. The method was validated with a small training data set of 109 binding cavities originating from a set of enzymes covering 12 different EC numbers. A second test data set of 1378 binding cavities, extracted from enzymes of 13 different EC numbers, was then used to prove the discriminating power of the algorithm and to demonstrate its applicability to large scale analyses. In all cases, members of the data set with the same EC number were placed into coherent regions on the map, with small distances between them. Different EC numbers are separated by large distances between the feature vectors. A third data set comprising three subfamilies of endopeptidases is used to demonstrate the ability of the algorithm to detect similar substructures between functionally related active sites. The algorithm can also be used to predict the function of novel proteins not considered in training data set.

Talk to us

Join us for a 30 min session where you can share your feedback and ask us any queries you have

Schedule a call

R Discovery Prime

R Discovery Prime

Large scale analysis of protein‐binding cavities using self‐organizing maps and wavelet‐based surface patches to describe functional properties, selectivity discrimination, and putative cross‐reactivity

Abstract

Talk to us

Similar Papers

More From: Proteins: Structure, Function, and Bioinformatics

Lead the way for us

Journal: Proteins: Structure, Function, and Bioinformatics	Publication Date: Nov 27, 2007
Citations: 21

Similar Papers

Retrieval, alignment, and clustering of computational models based on semantic annotations
Marvin Schulz ... Edda Klipp
Molecular Systems Biology | VOL. 7
Marvin Schulz, et. al.Marvin Schulz ... Edda Klipp
01 Jan 2010
Molecular Systems Biology | VOL. 7

The Edinburgh human metabolic network reconstruction and its functional analysis
Hongwu Ma ... Igor Goryanin
Molecular Systems Biology | VOL. 3
Hongwu Ma, et. al.Hongwu Ma ... Igor Goryanin
01 Jan 2007
Molecular Systems Biology | VOL. 3

Origin of heat capacity changes in a "nonclassical" hydrophobic interaction.
Neil R Syme ... Steve W Homans
ChemBioChem | VOL. 8
Neil R Syme, et. al.Neil R Syme ... Steve W Homans
12 Jul 2007
ChemBioChem | VOL. 8

Prediction of the Favorable Hydration Sites in a Protein Binding Pocket and Its Application to Scoring Function Formulation.
Yan Li ... Renxiao Wang
Journal of Chemical Information and Modeling | VOL. 60
Yan Li, et. al.Yan Li ... Renxiao Wang
13 May 2020
Journal of Chemical Information and Modeling | VOL. 60

Editage

Paperpal

R Discovery

Mind the Graph

R Discovery Prime

R Discovery Prime

Large scale analysis of protein‐binding cavities using self‐organizing maps and wavelet‐based surface patches to describe functional properties, selectivity discrimination, and putative cross‐reactivity

Abstract

Talk to us

Similar Papers

More From: Proteins: Structure, Function, and Bioinformatics