Folding non-homologous proteins by coupling deep-learning contact maps with I-TASSER assembly simulations
Folding non-homologous proteins by coupling deep-learning contact maps with I-TASSER assembly simulations
- Research Article
13
- 10.1186/1741-7007-11-44
- Apr 15, 2013
- BMC Biology
Segment assembly, structure alignment and iterative simulation in protein structure prediction
- Research Article
6
- 10.1096/fasebj.2020.34.s1.00169
- Apr 1, 2020
- The FASEB Journal
Protein structure prediction aims to determine the spatial location of every atom in protein molecules from the amino acid sequence by computational modeling. Depending on whether homologous structures are found in the Protein Data Bank (PDB), protein structure prediction have been historically categorized into template‐based modeling (TBM) and template‐free modeling (FM, or ab initio folding). In this talk, we first review recent progress in computer‐based protein structure prediction and show that the problem can be solved in principle (but not yet realistically) by TBM in case that fold‐recognition algorithms could identify the best structural templates from the PDB. Next, we discuss protein structure prediction results in the most recent community‐wide blind CASP experiments, showing that new approaches combining ab initio folding and deep neural‐network contact and distance predictions, which are built on residue coevolution data from multiple sequence alignments, can result in consistent and successful folding of proteins with sequence longer than 200–300 amino acids with a root‐mean‐square deviation below 3–4 Angstroms. This progress essentially breaks through the 50‐years‐old modeling border between TBM and FM and makes the success of high‐resolution structure prediction no longer relying on the PDB library. It also demonstrates the promise to solve the problem of protein structure prediction in a foreseeable future, by integrating deep machine learning techniques and the rapid advancement of genome sequencing databases, with the aid of advanced structure assembly simulation algorithms.Support or Funding InformationThis work is supported in part by the National Institute of General Medical Sciences (GM083107, GM116960), the National Institute of Allergy and Infectious Diseases (AI134678), and the National Science Foundation (DBI1564756, IIS1901191).
- Research Article
6
- 10.1016/j.entcs.2014.06.014
- Jul 1, 2014
- Electronic Notes in Theoretical Computer Science
MASTERS: A General Sequence-based MultiAgent System for Protein TERtiary Structure Prediction
- Conference Article
10
- 10.1109/icccnt.2013.6726753
- Jul 1, 2013
Proteins are essential parts of our life and participate in virtually every process within a cell. The understanding of protein structures is vital to determine the function of a protein. Protein structure prediction (PSP) from amino acid sequence is one of the high focus problems in bioinformatics today. This is due to the fact that the biological function of the protein is determined by its three dimensional structure. Thus, protein structure prediction is a fundamental area of computational biology. Its importance is intensed by large amounts of sequence data coming from PDB (Protein Data Bank) and the fact that experimentally methods such as X-ray crystallography or Nuclear Magnetic Resonance (NMR)which are used to determining protein structures remains very expensive and time consuming. For minimizing the time, computational methods are used for protein folding and structure prediction problem. In this paper results of protein p53 are discussed.
- Research Article
126
- 10.1016/j.jbc.2021.100870
- Jun 11, 2021
- The Journal of Biological Chemistry
Since Anfinsen demonstrated that the information encoded in a protein’s amino acid sequence determines its structure in 1973, solving the protein structure prediction problem has been the Holy Grail of structural biology. The goal of protein structure prediction approaches is to utilize computational modeling to determine the spatial location of every atom in a protein molecule starting from only its amino acid sequence. Depending on whether homologous structures can be found in the Protein Data Bank (PDB), structure prediction methods have been historically categorized as template-based modeling (TBM) or template-free modeling (FM) approaches. Until recently, TBM has been the most reliable approach to predicting protein structures, and in the absence of reliable templates, the modeling accuracy sharply declines. Nevertheless, the results of the most recent community-wide assessment of protein structure prediction experiment (CASP14) have demonstrated that the protein structure prediction problem can be largely solved through the use of end-to-end deep machine learning techniques, where correct folds could be built for nearly all single-domain proteins without using the PDB templates. Critically, the model quality exhibited little correlation with the quality of available template structures, as well as the number of sequence homologs detected for a given target protein. Thus, the implementation of deep-learning techniques has essentially broken through the 50-year-old modeling border between TBM and FM approaches and has made the success of high-resolution structure prediction significantly less dependent on template availability in the PDB library.
- Dissertation
- 10.32469/10355/94100
- Aug 1, 2022
[EMBARGOED UNTIL 8/1/2023] Building the high-quality structure of a protein from its amino acid sequence has important applications in protein engineering and drug design. The problem of accurate protein three-dimensional structure prediction from its amino acid sequence has not been completely solved yet. In the past several decades, many successful applications emerged on studying the protein structure prediction in one, two, and three dimensions. Prediction tools developed in one dimension as the protein secondary structure predictors are mature and generally applicable to the study of other onedimensional structure property predictions(e.g. protein solvent accessibility, protein disorder prediction) as well. Prediction of protein structure in two dimensions, i.e. protein inter-residue contact/distance prediction has been significantly advanced by the coevolutionary analysis and the application of deep learning with reasonably high accuracy. As a result, protein structure prediction in the three-dimension, especially the ab initio protein tertiary structure prediction has been dramatically improved by the accurate protein contact/distance prediction and deep learning techniques in the field of computer vision and natural language processing. In this thesis, efforts have been made on studying the protein sequence-to-structure relationship in one, two, and three dimensions. Four major contributions are listed: (a) The fast and effective method for protein secondary structure prediction--TransPross has been developed by applying the 1D transformer network and the attention mechanism. (b) The factors that affect the performance of deep learning in protein contact/ distance prediction have been systematically investigated. (c) The protein realvalue inter-residue distance predictor with the deep residual convolutional network- DeepDist was developed. (d) A novel end-to-end protein structure refinement tool called ATOMRefine was proposed. All the methods above have stand-alone software tools that are freely released to the public.||[EMBARGOED UNTIL 8/1/2023] Building the high-quality structure of a protein from its amino acid sequence has important applications in protein engineering and drug design. The problem of accurate protein three-dimensional structure prediction from its amino acid sequence has not been completely solved yet. In the past several decades, many successful applications emerged on studying the protein structure prediction in one, two, and three dimensions. Prediction tools developed in one dimension as the protein secondary structure predictors are mature and generally applicable to the study of other onedimensional structure property predictions(e.g. protein solvent accessibility, protein disorder prediction) as well. Prediction of protein structure in two dimensions, i.e. protein inter-residue contact/distance prediction has been significantly advanced by the coevolutionary analysis and the application of deep learning with reasonably high accuracy. As a result, protein structure prediction in the three-dimension, especially the ab initio protein tertiary structure prediction has been dramatically improved by the accurate protein contact/distance prediction and deep learning techniques in the field of computer vision and natural language processing. In this thesis, efforts have been made on studying the protein sequence-to-structure relationship in one, two, and three dimensions. Four major contributions are listed: (a) The fast and effective method for protein secondary structure prediction--TransPross has been developed by applying the 1D transformer network and the attention mechanism. (b) The factors that affect the performance of deep learning in protein contact/ distance prediction have been systematically investigated. (c) The protein realvalue inter-residue distance predictor with the deep residual convolutional network- DeepDist was developed. (d) A novel end-to-end protein structure refinement tool called ATOMRefine was proposed. All the methods above have stand-alone software tools that are freely released to the public.
- Book Chapter
9
- 10.1007/978-3-662-45049-9_42
- Jan 1, 2014
The biological function of the protein is folded by their spatial structure decisions, and therefore the process of protein folding is one of the most challenging problems in the field of bioinformatics. Although many heuristic algorithms have been proposed to solve the protein structure prediction (PSP) problem. The existing algorithms are far from perfect since PSP is an NP-problem. In this paper, we proposed artificial bee colony algorithm on 3D AB off-lattice model to PSP problem. In order to improve the global convergence ability and convergence speed of ABC algorithm, we adopt the new search strategy by combining the global solution into the search equation. Experimental results illustrate that the suggested algorithm is effective when the algorithm is applied to the Fibonacci sequences and four real protein sequences in the Protein Data Bank.
- Research Article
5
- 10.3233/fi-2015-1155
- Jan 1, 2015
- Fundamenta Informaticae
The protein structure folding is one of the most challenging problems in the field of bioinformatics. The main problem of protein structure prediction in the 3D toy model is to find the lowest energy conformation. Although many heuristic algorithms have been proposed to solve the protein structure prediction (PSP) problem, the existing algorithms are far from perfect since PSP is an NP-problem. In this paper, we proposed an artificial bee colony (ABC) algorithm based on the toy model to solve PSP problem. In order to improve the global convergence ability and convergence speed of the ABC algorithm, we adopt a new search strategy by combining the global solution into the search equation. Experimental results illustrate that the suggested algorithm can get the lowest energy when the algorithm is applied to the Fibonacci sequences and to four real protein sequences which come from the Protein Data Bank (PDB). Compared with the results obtained by PSO, LPSO, PSO-TS, PGATS, our algorithm is more efficient.
- Research Article
46
- 10.1002/prot.23046
- Jun 2, 2011
- Proteins: Structure, Function, and Bioinformatics
Prediction of protein structures from sequences is a fundamental problem in computational biology. Algorithms that attempt to predict a structure from sequence primarily use two sources of information. The first source is physical in nature: proteins fold into their lowest energy state. Given an energy function that describes the interactions governing folding, a method for constructing models of protein structures, and the amino acid sequence of a protein of interest, the structure prediction problem becomes a search for the lowest energy structure. Evolution provides an orthogonal source of information: proteins of similar sequences have similar structure, and therefore proteins of known structure can guide modeling. The relatively successful Rosetta approach takes advantage of the first, but not the second source of information during model optimization. Following the classic work by Andrej Sali and colleagues, we develop a probabilistic approach to derive spatial restraints from proteins of known structure using advances in alignment technology and the growth in the number of structures in the Protein Data Bank. These restraints define a region of conformational space that is high-probability, given the template information, and we incorporate them into Rosetta's comparative modeling protocol. The combined approach performs considerably better on a benchmark based on previous CASP experiments. Incorporating evolutionary information into Rosetta is analogous to incorporating sparse experimental data: in both cases, the additional information eliminates large regions of conformational space and increases the probability that energy-based refinement will hone in on the deep energy minimum at the native state.
- Abstract
1
- 10.1063/4.0000914
- Sep 1, 2025
- Structural Dynamics
The 2024 Nobel Prize in Chemistry was awarded to David Baker, Demis Hassabis, and John Jumper for computational protein design and protein structure prediction. These remarkable achievements depended critically on open access to the atomic-level, experimentally- determined, three-dimensional (3D) structures of macromolecules archived in the Protein Data Bank (PDB).RCSB.org is the research-focused web portal of the RCSB Protein Data Bank (RCSB PDB) that provides open access >230,000 PDB structures alongside >1 million Computed Structure Models (CSMs) generated using AlphaFold2 (from AlphaFoldDB), and RoseTTAFold and AlphaFold2 (from ModelArchive). For the avoidance of doubt, experimentally-determined PDB structures and CSMs delivered on RCSB.org by the RCSB PDB are clearly identified as to their respective provenance and reliability.RCSB.org users can query, organize, visualize, analyze, and compare experimental PDB structures and CSMs side-by-side by utilizing powerful tools:Search: User queries can be applied to all PDB structures and CSMs; PDB structures only; and can exclude either PDB structures or CSMs from the search results. N.B.: To reduce the likelihood of confusion with experimentally-determined structures, CSMs are not included in the search unless the RCSB.org user “opts in.”.View and Organize Results: By default, search results are ordered based on a relevancy score. Results can be resorted using various criteria (e.g., listing experimental PDB structures first, global per-residue confidence score (pLDDT)).Explore Similar Proteins: "Group" summary pages and search results simplify exploration of PDB structures with the same UniProt ID or sequence similarity or were deposited as part of the same study.Explore Individual Structures: Structure Summary Pages offer details of either experimental PDB structures and CSMs.Assess quality: Analogous to the validation slider for experimental structures, all CSMs report global and local confidence levels as pLDDT scores.Visualize in 3D: View experimental PDB structures and CSMs in Mol*. Use the standalone Mol* 3D Viewer to upload single or multiple data files, align structures, and run Structure Motif Search.Download: From Structure Summary Pages, download atomic coordinates for the PDB structure (various formats) or those of a CSM in ModelCIF format hosted by the corresponding external archive.RCSB PDB will continue to develop resources to support exploration of experimentally-determined PDB structures alongside CSMs, scaling infrastructure and operations to accommodate growth in the PDB archive and increased availability of predicted structures.RCSB PDB offers a variety of resources to support graduate students, postdoctoral fellows, and researchers in CSM exploration, including virtual and in-person events, with related materials published at PDB-101 (PDB101.RCSB.org) and user guidance documentation published at RCSB.org.PDB-101 hosts an additional collection of materials focused on protein prediction and protein design.RCSB PDB Core Operations are funded by National Science Foundation (DBI-2321666), US Department of Energy (DE-SC0019749), and National Cancer Institute, National Institute of Allergy and Infectious Diseases, and National Institute of General Medical Sciences of the National Institutes of Health under grant R01GM157729.
- Research Article
17
- 10.1016/j.swevo.2020.100711
- Jun 16, 2020
- Swarm and Evolutionary Computation
A GPU-based hybrid jDE algorithm applied to the 3D-AB protein structure prediction
- Research Article
290
- 10.1016/j.csbj.2019.12.011
- Jan 1, 2020
- Computational and Structural Biotechnology Journal
Protein Structure Prediction is a central topic in Structural Bioinformatics. Since the ’60s statistical methods, followed by increasingly complex Machine Learning and recently Deep Learning methods, have been employed to predict protein structural information at various levels of detail. In this review, we briefly introduce the problem of protein structure prediction and essential elements of Deep Learning (such as Convolutional Neural Networks, Recurrent Neural Networks and basic feed-forward Neural Networks they are founded on), after which we discuss the evolution of predictive methods for one-dimensional and two-dimensional Protein Structure Annotations, from the simple statistical methods of the early days, to the computationally intensive highly-sophisticated Deep Learning algorithms of the last decade. In the process, we review the growth of the databases these algorithms are based on, and how this has impacted our ability to leverage knowledge about evolution and co-evolution to achieve improved predictions. We conclude this review outlining the current role of Deep Learning techniques within the wider pipelines to predict protein structures and trying to anticipate what challenges and opportunities may arise next.
- Book Chapter
9
- 10.1007/978-3-642-13214-8_30
- Jan 1, 2010
This paper proposes an approach to the protein structure prediction (PSP) problem that inserts solutions provided by template-modeling procedures as individuals in the population of a multi-objective evolutionary algorithm. This way, our procedure represents a hybrid approach that takes advantage of previous knowledge about the known protein structures to improve the effectiveness of an ab initio procedure for the PSP problem. Moreover, the procedure benefits from a parallel and distributed implementation that allows faster and wider exploration of the conformation space. The experimental results obtained from the present implementation of our procedure show improvements with respect to previously proposed procedures in the proteins selected as benchmarks from the CASP8 set (up to 28% of RMSD improvement with respect to TASSER).
- Conference Article
20
- 10.1109/cec.2011.5949957
- Jun 1, 2011
One of the main research problems in Structural Bioinformatics is related to the prediction of three-dimensional structures (3-D) of polypeptides or proteins. The rate at which amino acid sequences are identified is increasing faster than the 3-D protein structure determination by experimental methods. Computational prediction methods have been developed during the last years, but the problem still remains challenging because of the complexity and high dimensionality of a protein conformational search space. In this article we present a hybrid genetic algorithm for the Protein Structure Prediction (PSP) Problem. A genetic algorithm is combined with a structured population, and it is hybridized with a path-relinking procedure that helps the algorithm to scape from local minima. We perform a set of experiments and show that the proposed hybrid genetic algorithm is effective in finding good quality solutions for the PSP Problem.
- Research Article
2
- 10.1016/j.eswa.2012.08.003
- Sep 7, 2012
- Expert Systems With Applications
A molecular dynamics and knowledge-based computational strategy to predict native-like structures of polypeptides