Improving the Physical Realism and Structural Accuracy of Protein Models by a Two-Step Atomic-Level Energy Minimization
Improving the Physical Realism and Structural Accuracy of Protein Models by a Two-Step Atomic-Level Energy Minimization
- Research Article
23
- 10.1093/bioinformatics/btad298
- May 4, 2023
- Bioinformatics
The state-of-art protein structure prediction methods such as AlphaFold are being widely used to predict structures of uncharacterized proteins in biomedical research. There is a significant need to further improve the quality and nativeness of the predicted structures to enhance their usability. In this work, we develop ATOMRefine, a deep learning-based, end-to-end, all-atom protein structural model refinement method. It uses a SE(3)-equivariant graph transformer network to directly refine protein atomic coordinates in a predicted tertiary structure represented as a molecular graph. The method is first trained and tested on the structural models in AlphaFoldDB whose experimental structures are known, and then blindly tested on 69 CASP14 regular targets and 7 CASP14 refinement targets. ATOMRefine improves the quality of both backbone atoms and all-atom conformation of the initial structural models generated by AlphaFold. It also performs better than two state-of-the-art refinement methods in multiple evaluation metrics including an all-atom model quality score-the MolProbity score based on the analysis of all-atom contacts, bond length, atom clashes, torsion angles, and side-chain rotamers. As ATOMRefine can refine a protein structure quickly, it provides a viable, fast solution for improving protein geometry and fixing structural errors of predicted structures through direct coordinate refinement. The source code of ATOMRefine is available in the GitHub repository (https://github.com/BioinfoMachineLearning/ATOMRefine). All the required data for training and testing are available at https://doi.org/10.5281/zenodo.6944368.
- Research Article
203
- 10.1002/prot.24167
- Sep 26, 2012
- Proteins: Structure, Function, and Bioinformatics
One of the major limitations of computational protein structure prediction is the deviation of predicted models from their experimentally derived true, native structures. The limitations often hinder the possibility of applying computational protein structure prediction methods in biochemical assignment and drug design that are very sensitive to structural details. Refinement of these low-resolution predicted models to high-resolution structures close to the native state, however, has proven to be extremely challenging. Thus, protein structure refinement remains a largely unsolved problem. Critical assessment of techniques for protein structure prediction (CASP) specifically indicated that most predictors participating in the refinement category still did not consistently improve model quality. Here, we propose a two-step refinement protocol, called 3Drefine, to consistently bring the initial model closer to the native structure. The first step is based on optimization of hydrogen bonding (HB) network and the second step applies atomic-level energy minimization on the optimized model using a composite physics and knowledge-based force fields. The approach has been evaluated on the CASP benchmark data and it exhibits consistent improvement over the initial structure in both global and local structural quality measures. 3Drefine method is also computationally inexpensive, consuming only few minutes of CPU time to refine a protein of typical length (300 residues). 3Drefine web server is freely available at http://sysbio.rnet.missouri.edu/3Drefine/.
- Research Article
53
- 10.1371/journal.pone.0069648
- Jul 19, 2013
- PLoS ONE
Protein structure refinement refers to the process of improving the qualities of protein structures during structure modeling processes to bring them closer to their native states. Structure refinement has been drawing increasing attention in the community-wide Critical Assessment of techniques for Protein Structure prediction (CASP) experiments since its addition in 8th CASP experiment. During the 9th and recently concluded 10th CASP experiments, a consistent growth in number of refinement targets and participating groups has been witnessed. Yet, protein structure refinement still remains a largely unsolved problem with majority of participating groups in CASP refinement category failed to consistently improve the quality of structures issued for refinement. In order to alleviate this need, we developed a completely automated and computationally efficient protein 3D structure refinement method, i3Drefine, based on an iterative and highly convergent energy minimization algorithm with a powerful all-atom composite physics and knowledge-based force fields and hydrogen bonding (HB) network optimization technique. In the recent community-wide blind experiment, CASP10, i3Drefine (as ‘MULTICOM-CONSTRUCT’) was ranked as the best method in the server section as per the official assessment of CASP10 experiment. Here we provide the community with free access to i3Drefine software and systematically analyse the performance of i3Drefine in strict blind mode on the refinement targets issued in CASP10 refinement category and compare with other state-of-the-art refinement methods participating in CASP10. Our analysis demonstrates that i3Drefine is only fully-automated server participating in CASP10 exhibiting consistent improvement over the initial structures in both global and local structural quality metrics. Executable version of i3Drefine is freely available at http://protein.rnet.missouri.edu/i3drefine/.
- Research Article
13
- 10.1186/1741-7007-11-44
- Apr 15, 2013
- BMC Biology
Segment assembly, structure alignment and iterative simulation in protein structure prediction
- Research Article
18
- 10.1021/jp209657n
- Feb 16, 2012
- The Journal of Physical Chemistry B
The challenge in protein structure prediction using homology modeling is the lack of reliable methods to refine the low resolution homology models. Unconstrained all-atom molecular dynamics (MD) does not serve well for structure refinement due to its limited conformational search. We have developed and tested the constrained MD method, based on the generalized Newton-Euler inverse mass operator (GNEIMO) algorithm for protein structure refinement. In this method, the high-frequency degrees of freedom are replaced with hard holonomic constraints and a protein is modeled as a collection of rigid body clusters connected by flexible torsional hinges. This allows larger integration time steps and enhances the conformational search space. In this work, we have demonstrated the use of torsional GNEIMO method without using any experimental data as constraints, for protein structure refinement starting from low-resolution decoy sets derived from homology methods. In the eight proteins with three decoys for each, we observed an improvement of ~2 Å in the rmsd in coordinates to the known experimental structures of these proteins. The GNEIMO trajectories also showed enrichment in the population density of native-like conformations. In addition, we demonstrated structural refinement using a "freeze and thaw" clustering scheme with the GNEIMO framework as a viable tool for enhancing localized conformational search. We have derived a robust protocol based on the GNEIMO replica exchange method for protein structure refinement that can be readily extended to other proteins and possibly applicable for high throughput protein structure refinement.
- Abstract
10
- 10.1016/j.bpj.2011.11.160
- Jan 1, 2012
- Biophysical Journal
Structure Refinement of Protein Low Resolution Models using GNEIMO Constrained Dynamics Method
- Research Article
193
- 10.1110/ps.03381404
- Jan 1, 2004
- Protein Science
The use of classical molecular dynamics simulations, performed in explicit water, for the refinement of structural models of proteins generated ab initio or based on homology has been investigated. The study involved a test set of 15 proteins that were previously used by Baker and coworkers to assess the efficiency of the ROSETTA method for ab initio protein structure prediction. For each protein, four models generated using the ROSETTA procedure were simulated for periods of between 5 and 400 nsec in explicit solvent, under identical conditions. In addition, the experimentally determined structure and the experimentally derived structure in which the side chains of all residues had been deleted and then regenerated using the WHATIF program were simulated and used as controls. A significant improvement in the deviation of the model structures from the experimentally determined structures was observed in several cases. In addition, it was found that in certain cases in which the experimental structure deviated rapidly from the initial structure in the simulations, indicating internal strain, the structures were more stable after regenerating the side-chain positions. Overall, the results indicate that molecular dynamics simulations on a tens to hundreds of nanoseconds time scale are useful for the refinement of homology or ab initio models of small to medium-size proteins.
- Research Article
24
- 10.2174/092986612803217015
- Sep 1, 2012
- Protein & Peptide Letters
The gap between known protein sequences and structures is increasing rapidly and experimental methods alone will not be able to fill in this gap. Therefore it is necessary to use computational methods to predict protein structures. Template based modeling methods could be used for sequences, which have detectable relationship with sequences of one or more experimentally determined protein structures. For predicting the structure of proteins, which does not share a detectable sequence relationship with experimental structures, ab initio protein structure prediction techniques must be used. The methods under ab initio protein structure prediction category aim to predict the structure of a protein from the sequence information alone, without any explicit use of previously known structures. These methods use thermodynamic principles and try to identify the native structure of a protein as the global minimum of a potential energy landscape. However, such methods are computationally complex and are extraordinarily challenging. There has been significant progress in the development of ab inito protein structure prediction methods over the past few years. This review describes the basic principles, the complexity, challenges and recent progresses of ab initio protein structure prediction.
- Research Article
- 10.1109/tcbb.2024.3422288
- Mar 1, 2025
- IEEE transactions on computational biology and bioinformatics
The goal of protein structure refinement is to enhance the precision of predicted protein models, particularly at the residue level of the local structure. Existing refinement approaches primarily rely on physics, whereas molecular simulation methods are resource-intensive and time-consuming. In this study, we employ deep learning methods to extract structural constraints from protein structure residues to assist in protein structure refinement. We introduce a novel method, AnglesRefine, which focuses on a protein's secondary structure and employs transformer to refine various protein structure angles (psi, phi, omega, CA_C_N_angle, C_N_CA_angle, N_CA_C_angle), ultimately generating a superior protein model based on the refined angles. We evaluate our approach against other cutting-edge methods using the CASP11-14 and CASP15 datasets. Experimental outcomes indicate that our method generally surpasses other techniques on the CASP11-14 test dataset, while performing comparably or marginally better on the CASP15 test dataset. Our method consistently demonstrates the least likelihood of model quality degradation, e.g., the degradation percentage of our method is less than 10%, while other methods are about 50%. Furthermore, as our approach eliminates the need for conformational search and sampling, it significantly reduces computational time compared to existing refinement methods.
- Front Matter
8
- 10.1186/gb-2000-1-1-comment002
- Jan 1, 2000
- Genome Biology
The Grail problem
- Research Article
139
- 10.1002/prot.21345
- May 1, 2007
- Proteins: Structure, Function, and Bioinformatics
Recent advances in efficient and accurate treatment of solvent with the generalized Born approximation (GB) have made it possible to substantially refine the protein structures generated by various prediction tools through detailed molecular dynamics simulations. As demonstrated in a recent CASPR experiment, improvement can be quite reliably achieved when the initial models are sufficiently close to the native basin (e.g., 3-4 A C(alpha) RMSD). A key element to effective refinement is to incorporate reliable structural information into the simulation protocol. Without intimate knowledge of the target and prediction protocol used to generate the initial structural models, it can be assumed that the regular secondary structure elements (helices and strands) and overall fold topology are largely correct to start with, such that the protocol limits itself to the scope of refinement and focuses the sampling in vicinity of the initial structure. The secondary structures can be enforced by dihedral restraints and the topology through structural contacts, implemented as either multiple pair-wise C(alpha) distance restraints or a single sidechain distance matrix restraint. The restraints are weakly imposed with flat-bottom potentials to allow sufficient flexibility for structural rearrangement. Refinement is further facilitated by enhanced sampling of advanced techniques such as the replica exchange method (REX). In general, for single domain proteins of small to medium sizes, 3-5 nanoseconds of REX/GB refinement simulations appear to be sufficient for reasonable convergence. Clustering of the resulting structural ensembles can yield refined models over 1.0 A closer to the native structure in C(alpha) RMSD. Substantial improvement of sidechain contacts and rotamer states can also be achieved in most cases. Additional improvement is possible with longer sampling and knowledge of the robust structural features in the initial models for a given prediction protocol. Nevertheless, limitations still exist in sampling as well as force field accuracy, manifested as difficulty in refinement of long and flexible loops.
- Research Article
25
- 10.1002/jcc.21040
- Jul 9, 2008
- Journal of Computational Chemistry
Tetracyclines (Tcs) are an important family of antibiotics that bind to the ribosome and several proteins. To model Tc interactions with protein and RNA, we have developed a molecular mechanics force field for 12 tetracyclines, consistent with the CHARMM force field. We considered each Tc variant in its zwitterionic tautomer, with and without a bound Mg(2+). We used structures from the Cambridge Crystallographic Data Base to identify the conformations likely to be present in solution and in biomolecular complexes. A conformational search by simulated annealing was undertaken, using the MM3 force field, for tetracycline, anhydrotetracycline, doxycycline, and tigecycline. Resulting, low-energy structures were optimized with an ab initio method. We found that Tc and its analogs all adopt an extended conformation in the zwitterionic tautomer and a twisted one in the neutral tautomer, and the zwitterionic-extended state is the most stable in solution. Intermolecular force field parameters were derived from a standard supermolecule approach: we considered the ab initio energies and geometries of a water molecule interacting with each Tc analog at several different positions. The final, rms deviation between the ab initio and force field energies, averaged over all forms, was 0.35 kcal/mol. Intramolecular parameters were adopted from either the standard CHARMM force field, the ab initio structure, or the earlier, plain Tc force field. The model reproduces the ab initio geometry and flexibility of each Tc. As tests, we describe MD and free energy simulations of a solvated complex between three Tcs and the Tet repressor protein.
- Research Article
72
- 10.1111/j.1432-1033.1992.tb16972.x
- Jun 1, 1992
- European Journal of Biochemistry
546 NOESY cross-peak volumes were measured in the two-dimensional NOESY spectrum of proteolytic fragment 163-231 of bacterioopsin in organic solution. These data and 42 detected hydrogen bonds were applied for determining the peptide spatial structure. The fold of the polypeptide chain was determined by local structure analysis, a distance geometry approach and systematic search for energetically allowed side-chain rotamers which are consistent with experimental NOESY cross-peak volumes. The effective rotational correlation time of 6 ns for the molecule was evaluated from optimization of the local structure to meet NOE data and from the dependence on mixing time of the NiH/Ci alpha H cross-peak volumes of the residues in alpha-helical conformation. The resulting structure has two well defined alpha-helical regions, 168-191 and 198-227, with root-mean-square deviation 44 pm and 69 pm, respectively, between the backbone atoms in 14 final energy refined conformations. The alpha-helices correspond to transmembrane segments F and G of bacteriorhodopsin. The segment F contains proline 186, which introduces a kink of about 25 degrees with a disruption of the hydrogen bond with the NH group of the following residue. The segments are connected by a flexible loop region 192-197. Torsion angles chi 1 are unequivocally defined for 62% of side chains in the alpha-helices but half of them differ from electron cryo-microscopy (ECM) model of bacteriorhodopsin, apparently because of the low resolution of ECM. Nevertheless, the F and G segments can be packed as in the ECM model and with side-chain conformations consistent with all NMR data in solution.
- Research Article
19
- 10.1016/j.sbi.2013.01.010
- Apr 1, 2013
- Current opinion in structural biology
Computational methods for high resolution prediction and refinement of protein structures.
- Research Article
74
- 10.1021/ja971461w
- Dec 1, 1997
- Journal of the American Chemical Society
We have investigated the carbon-13 solution nuclear magnetic resonance (NMR) chemical shifts of Cα, Cβ, and Cγ carbons of 19 valine residues in a vertebrate calmodulin, a nuclease from Staphylococcus aureus, and a ubiquitin. Using empirical chemical shift surfaces to predict Cα, Cβ shifts from known, X-ray φ,ψ values, we find moderate accord between prediction and experiment. Ab initio calculations with coupled Hartree−Fock (HF) methods and X-ray structures yield poor agreement with experiment. There is an improvement in the ab initio results when the side chain χ1 torsion angles are adjusted to their lowest energy conformers, using either ab initio quantum chemical or empirical methods, and a further small improvement when the effects of peptide-backbone charge fields are introduced. However, although the theoretical and experimental results are highly correlated (R2 ∼ 0.90), the observed slopes of ∼−0.6−0.8 are less than the ideal value of −1, even when large uniform basis sets are used. Use of density functional theory (DFT) methods improves the quality of the predictions for both Cα (slope = −1.1, R2 = 0.91) and Cβ (slope = −0.93, R2 = 0.89), as well as giving moderately good results for Cγ. This effect is thought to arise from a small, conformationally-sensitive contribution to shielding arising from electron correlation. Additional shielding calculations on model compounds reveal similar effects. Results for valine residues in interleukin-1β are less highly correlated, possibly due to larger crystal−solution structural differences. When taken together, these results for 19 valine residues in 3 proteins indicate that choosing the lowest energy χ1 conformer together with X-ray φ,ψ values enables the successful prediction of both Cα and Cβ shifts, with DFT giving close to ideal slopes and R2 values between theory and experiment. These results strongly suggest that the most highly populated valine side-chain conformers are those having the lowest (computationally determined) energy, as evidenced by the ability to predict essentially all Cα, Cβ chemical shifts in calmodulin, SNase, and ubiquitin, as well as moderate accord for Cγ. These observations suggest a role for chemical shifts and energy minimization/geometry optimization in the refinement of protein structures in solution, and potentially in the solid state as well.