AI-based quality assessment methods for protein structure models from cryo-EM.
AI-based quality assessment methods for protein structure models from cryo-EM.
- Research Article
- 10.1107/s2053273323099473
- Jul 7, 2023
- Acta Crystallographica Section A Foundations and Advances
In recent years, an increasing number of protein and nucleotide structures have been modeled from cryo-electron microscopy (cryo-EM) maps. However, even though the EM map resolution has generally improved steadily over the past years, there are still many situations where modeling errors occur in high-resolution EM maps, or modelers face difficulties in modeling biomolecular structures due to locally low resolution in the map. To address such challenges, we have applied deep learning to three tasks: model quality assessment, protein structure modeling, and DNA/RNA structure modeling in cryo-EM maps. 1: Model Quality Assessment Modeling a protein structure into a cryo-EM map is a challenging task. One of the main difficulties is assigning the correct amino acids to their corresponding positions. Moreover, even with high-quality maps, there is always a risk of human error in the modeling process. To ensure the resulting atomic model is as accurate as possible, it's essential to perform rigorous validation using appropriate methods. To validate protein structure models in cryo-EM maps, our group developed a novel method based on the Deep-learning-based Amino-acid-wise model Quality (DAQ) score. In the DAQ score, the neural network detects specific map features for protein amino acid residue types, Cα atoms, and secondary structures, and computes the likelihood that each residue assignment is correct. By quantifying the incompatibilities between the protein model and the EM map at the amino acid level, the DAQ score provides a more accurate and sensitive measure of model quality compared to other methods [1] . Overall, the DAQ score offers a powerful tool for assessing protein structure models in EM maps and advancing cryo-EM research. The DAQ score can be computed on the Google Colab site (https://bit.ly/daqscore) or local machine by installing the code from (https://github.com/kiharalab/DAQ). Our group has also recently released the DAQ-Score Database [2] (https://daqdb.kiharalab.org/), which provides precomputed quality assessment results for protein models deposited in the Protein Data Bank (PDB) and their corresponding cryo-EM maps in the Electron Microscopy Data Bank (EMDB). Currently, the DAQ-Score Database contains over 152,129 protein chain models from 9,469 PDB entries derived from cryo-EM maps. In addition, the database provides the DAQ scores for multiple previous major versions of models if they exist. An example of a database entry is shown in Figure 1 , which shows the first and revised version of the model for PDB ID: 7JSN Chain B (ID: 22458_7jsn_B_v1-1 and 22458_7jsn_B_v2-0). The computed DAQ scores are presented in a color code on a model within an interactive structure viewer, coupled with a graph showing the three DAQ score types along the residue sequence number. Figure 1 shows that the first model version (top) has regions that indicate negative DAQ scores (i.e., low quality). The corresponding regions in the revised model (bottom) were updated to positive DAQ scores, indicating substantial improvement 2: DeepMainmast: Protein structure modeling Protein structure modeling from a cryo-EM map is challenging, particularly when the resolution is worse than about 3 Å. To address this problem, we have developed an integrated protein structure modeling protocol called DeepMainmast [3]. This protocol employs a new de novo protein main-chain tracing method that uses deep learning to identify positions of Cα atoms and the types of amino acids. The core process of DeepMainmast employs an effective main-chain tracing approach, the Vehicle Routing Problem solver, and Constraint Problem Solver. Additionally, the protocol can accurately assign chain identity to the structure models of homo-multimers. To enhance the performance of the protocol, we also incorporate AlphaFold2 models when applicable. These models provide valuable information that can improve the accuracy of the resulting protein structures. Overall, DeepMainmast is a powerful tool that can help researchers overcome the challenges of protein structure modeling from cryo-EM maps. Our benchmarking results demonstrate that DeepMainmast substantially outperforms existing methods on the benchmark dataset. Compared to AlphaFold2, DeepMainmast achieves higher accuracy on a larger number of maps within the dataset, which consists of 178 high-resolution maps. Figure 2 shows an example of the modeling result for EMD-6551. DeepMainmast generated the accurate model with the correct chain ID assignment. The code is available at https://github.com/kiharalab/DeepMainMast. CryoREAD: DNA/RNA structure modeling Modelling the structure of DNA/RNA from cryo-EM maps is generally more challenging than protein structure modelling due to a number of factors. For example, DNA and RNA molecules can exhibit greater flexibility and variability in their structure than proteins, and available 3D structure data for DNA/RNA is significantly less than that of proteins. As a result, most biomolecular structure modeling software is primarily designed for proteins. To overcome this challenge, we have developed CryoREAD [4] , which is a novel method for automated de novo DNA/RNA structure modeling from cryo-EM maps of a resolution range, between 2.0 Å to 5.0 Å. The method uses a deep neural network to identify the potential positions of phosphate, sugar, and bases, construct the backbone structure, map the nucleic acid sequence along the backbone, and construct a full atom model. Figure 3 illustrates the input EM map (EMD-12217), outputs of deep learning, and the final structure model with the native structure (PDB-ID:7BL4). Based on our benchmarking of 68 cryo-EM maps, on average, 84.9% of the atoms were correctly placed within a 5 Å, and 52.1% of nucleotides were correctly identified. The CryoREAD is available at https://github.com/kiharalab/CryoREAD.
- Research Article
245
- 10.1016/j.matt.2020.10.021
- Nov 9, 2020
- Matter
Cathode-Electrolyte Interphase in Lithium Batteries Revealed by Cryogenic Electron Microscopy
- Research Article
16
- 10.1007/978-981-10-1503-8_3
- Jan 1, 2016
- Advances in experimental medicine and biology
Protein structure prediction and modeling provide a tool for understanding protein functions by computationally constructing protein structures from amino acid sequences and analyzing them. With help from protein prediction tools and web servers, users can obtain the three-dimensional protein structure models and gain knowledge of functions from the proteins. In this chapter, we will provide several examples of such studies. As an example, structure modeling methods were used to investigate the relation between mutation-caused misfolding of protein and human diseases including epilepsy and leukemia. Protein structure prediction and modeling were also applied in nucleotide-gated channels and their interaction interfaces to investigate their roles in brain and heart cells. In molecular mechanism studies of plants, rice salinity tolerance mechanism was studied via structure modeling on crucial proteins identified by systems biology analysis; trait-associated protein-protein interactions were modeled, which sheds some light on the roles of mutations in soybean oil/protein content. In the age of precision medicine, we believe protein structure prediction and modeling will play more and more important roles in investigating biomedical mechanism of diseases and drug design.
- Research Article
4
- 10.1101/2023.10.19.563151
- Nov 21, 2023
- bioRxiv
Structure modeling from maps is an indispensable step for studying proteins and their complexes with cryogenic electron microscopy (cryo-EM). Although the resolution of determined cryo-EM maps has generally improved, there are still many cases where tracing protein main-chains is difficult, even in maps determined at a near atomic resolution. Here, we have developed a protein structure modeling method, called DeepMainmast, which employs deep learning to capture the local map features of amino acids and atoms to assist main-chain tracing. Moreover, since Alphafold2 demonstrates high accuracy in protein structure prediction, we have integrated complementary strengths of de novo density tracing using deep learning with Alphafold2’s structure modeling to achieve even higher accuracy than each method alone. Additionally, the protocol is able to accurately assign chain identity to the structure models of homo-multimers.
- Research Article
49
- 10.1038/s41592-023-02032-5
- Oct 2, 2023
- Nature methods
DNA and RNA play fundamental roles in various cellular processes, where their three-dimensional structures provide information critical to understanding the molecular mechanisms of their functions. Although an increasing number of nucleic acid structures and their complexes with proteins are determined by cryogenic electron microscopy (cryo-EM), structure modeling for DNA and RNA remains challenging particularly when the map is determined at a resolution coarser than atomic level. Moreover, computational methods for nucleic acid structure modeling are relatively scarce. Here, we present CryoREAD, a fully automated de novo DNA/RNA atomic structure modeling method using deep learning. CryoREAD identifies phosphate, sugar and base positions in a cryo-EM map using deep learning, which are traced and modeled into a three-dimensional structure. When tested on cryo-EM maps determined at 2.0 to 5.0 Å resolution, CryoREAD built substantially more accurate models than existing methods. We also applied the method to cryo-EM maps of biomolecular complexes in severe acute respiratory syndrome coronavirus 2 (SARS-CoV-2).
- Research Article
6
- 10.1186/s12900-018-0097-0
- Dec 1, 2018
- BMC Structural Biology
BackgroundIn the backdrop of challenge to obtain a protein structure under the known limitations of both experimental and theoretical techniques, the need of a fast as well as accurate protein structure evaluation method still exists to substantially reduce a huge gap between number of known sequences and structures. Among currently practiced theoretical techniques, homology modelling backed by molecular dynamics based optimization appears to be the most popular one. However it suffers from contradictory indications of different validation parameters generated from a set of protein models which are predicted against a particular target protein. For example, in one model Ramachandran Score may be quite high making it acceptable, whereas, its potential energy may not be very low making it unacceptable and vice versa. Towards resolving this problem, the main objective of this study was fixed as to utilize a simple experimentally derived output, Surface Roughness Index of concerned protein of unknown structure as an intervening agent that could be obtained using ordinary microscopic images of heat denatured aggregates of the same protein.ResultIt was intriguing to observe that direct experimental knowledge of the concerned protein, however simple it may be, might give insight on acceptability of its particular structural model out of a confusion set of models generated from database driven comparative technique for structure prediction. The result obtained from a widely varying structural class of proteins indicated that speed of protein structure evaluation can be further enhanced without compromising with accuracy by recruiting simple experimental output.ConclusionIn this work, a semi-empirical methodological approach was provided for improving protein structure evaluation. It showed that, once structure models of a protein were obtained through homology technique, the problem of selection of a best model out of a confusion set of Pareto-optimal structures could be resolved by employing a structure agent directly obtainable through experiment with the same protein as experimental ingredient. Overall, in the backdrop of getting a reasonably accurate protein structure of pathogens causing epidemics or biological warfare, such approach could be of use as a plausible solution for fast drug design.
- Research Article
23
- 10.3389/fphys.2019.01452
- Dec 3, 2019
- Frontiers in Physiology
Despite significant effort on understanding complex biological systems, we lack a unified theory for modeling, inference, analysis, and efficient control of their dynamics in uncertain environments. These problems are made even more challenging when considering that only limited and noisy information is accessible for modeling, which can prove insufficient for explaining, and predicting the behavior of complex systems. For instance, missing information hampers the capabilities of analytical tools to uncover the true degrees of freedom and infer the model structure and parameters of complex biological systems. Toward this end, in this paper, we discuss several important mathematical challenges that could open new theoretical avenues in studying complex systems: (1) By understanding the universal laws characterizing the asymmetric statistics of magnitude increments and the complex space-time interdependency within one process and across many processes, we can develop a class of compact yet accurate mathematical models capable to potentially providing higher degree of predictability, and more efficient control strategies. (2) In order to better predict the onset of disease and their root cause, as well as potentially discover more efficient quality-of-life (QoL)-control strategies, we need to develop mathematical strategies that not only are capable to discover causal interactions and their corresponding mathematical expressions for space and time operators acting on biological processes, but also mathematical and algorithmic techniques to identify the number of unknown unknowns (UUs) and their interdependency with the observed variables. (3) Lastly, to improve the QoL of control strategies when facing intra- and inter-patient variability, the focus should not only be on specific values and ranges for biological processes, but also on optimizing/controlling knob variables that enforce a specific spatiotemporal multifractal behavior that corresponds to an initial healthy (patient specific) behavior. All in all, the modeling, analysis and control of complex biological collective systems requires a deeper understanding of the multifractal properties of high dimensional heterogeneous and noisy data streams and new algorithmic tools that exploit geometric, statistical physics, and information theoretic concepts to deal with these data challenges.
- Research Article
47
- 10.1186/1472-6807-7-43
- Jun 29, 2007
- BMC Structural Biology
BackgroundAlthough experimental methods for determining protein structure are providing high resolution structures, they cannot keep the pace at which amino acid sequences are resolved on the scale of entire genomes. For a considerable fraction of proteins whose structures will not be determined experimentally, computational methods can provide valuable information. The value of structural models in biological research depends critically on their quality. Development of high-accuracy computational methods that reliably generate near-experimental quality structural models is an important, unsolved problem in the protein structure modeling.ResultsLarge sets of structural decoys have been generated using reduced conformational space protein modeling tool CABS. Subsequently, the reduced models were subject to all-atom reconstruction. Then, the resulting detailed models were energy-minimized using state-of-the-art all-atom force field, assuming fixed positions of the alpha carbons. It has been shown that a very short minimization leads to the proper ranking of the quality of the models (distance from the native structure), when the all-atom energy is used as the ranking criterion. Additionally, we performed test on medium and low accuracy decoys built via classical methods of comparative modeling. The test placed our model evaluation procedure among the state-of-the-art protein model assessment methods.ConclusionThese test computations show that a large scale high resolution protein structure prediction is possible, not only for small but also for large protein domains, and that it should be based on a hierarchical approach to the modeling protocol. We employed Molecular Mechanics with fixed alpha carbons to rank-order the all-atom models built on the scaffolds of the reduced models. Our tests show that a physic-based approach, usually considered computationally too demanding for large-scale applications, can be effectively used in such studies.
- Research Article
3
- 10.1101/2024.09.06.611715
- Sep 11, 2024
- bioRxiv
Motivation:Cryogenic Electron Microscopy (cryo-EM) is a core experimental technique used to determine the structure of macromolecules such as proteins. However, the effectiveness of cryo-EM is often hindered by the noise and missing density values in cryo-EM density maps caused by experimental conditions such as low contrast and conformational heterogeneity. Although various global and local map sharpening techniques are widely employed to improve cryo-EM density maps, it is still challenging to efficiently improve their quality for building better protein structures from them.Results:In this study, we introduce CryoTEN - a three-dimensional U-Net style transformer to improve cryo-EM maps effectively. CryoTEN is trained using a diverse set of 1,295 cryo-EM maps as inputs and their corresponding simulated maps generated from known protein structures as targets. An independent test set containing 150 maps is used to evaluate CryoTEN, and the results demonstrate that it can robustly enhance the quality of cryo-EM density maps. In addition, the automatic de novo protein structure modeling shows that the protein structures built from the density maps processed by CryoTEN have substantially better quality than those built from the original maps. Compared to the existing state-of-the-art deep learning methods for enhancing cryo-EM density maps, CryoTEN ranks second in improving the quality of density maps, while running > 10 times faster and requiring much less GPU memory than them.Availability and implementation:The source code and data is freely available at https://github.com/jianlin-cheng/cryoten
- Supplementary Content
- 10.7907/p18t-5j69.
- Jun 6, 2020
Cells within biological systems are constantly adjusting their protein synthesis in response to various environmental changes. To study the rapid cellular regulations in complex biological systems, global proteomic profiling provides important information on system-level regulations, yet physiological properties characteristic of individual cellular subpopulations could be hidden under the characterization. Instead, cell-selective proteomic profiling allows researchers to reveal the heterogeneities in biological systems with phenotypically and even genetically distinct subpopulations under different microenvironments. Chapter 1 describes the development of bioorthogonal noncanonical amino acid tagging (BONCAT) for proteomic profiling with resolution in both space and time: its initial role is protein labeling with temporal resolution via pulse-addition of noncanonical amino acid, which could be recognized by endogenous aminoacyl tRNA-synthetase (aaRS), into systems of interest; later on, mutant aaRSs are identified through mutant synthetase library screening, which allows for efficient incorporation of various types of noncanonical amino acids that could hardly be activated by endogenous machineries. The identification and exploitation of mutant aaRSs allow sensitive cellular selectivity during protein labeling. With unprecedented spatiotemporal resolution of BONCAT, and the advancement in high-resolution mass spectrometry and computational algorithms, BONCAT is a powerful technique for selective proteomic profiling to study physiological regulations in a wide range of complex biological systems. Chapter 2 describes the application of the BONCAT method in cell-selective proteomic profiling in Pseudomonas aeruginosa biofilms. In this work, we targeted an iron-starved subpopulation in biofilms and compared its proteomic profile with that of the entire system. Key gene and pathway regulations in the subpopulation are found through the analysis of the proteomic data, which suggest that iron-starved cells shift their priority towards housing keeping pathways, adapt an energy- and resources-saving mode to cope with their harsh local environmental conditions, and get prepared to disperse for better survival. Analysis of poorly studied proteins highly upregulated in the subpopulation led to the discovery of a previously uncharacterized protein (PA14_52000) that is potentially related to iron acquisition. The transposon insertion mutant PA14_52000::tn showed significantly enhanced pyoverdine production in rich medium and reduced biofilm formation. Chapter 3 describes the study of physiological regulations in Bacillus subtilis K-state subpopulation via BONCAT. A subset of B. subtilis cells, typically 10% - 20% of the entire population, enter K-state in a stochastic manner. With the low level of K-state entry rate and high randomness, we challenged BONCAT to specifically capture gene and pathway regulations in K-state cells and compared the proteomic profiling with that of the entire population. Regardless of the difficulties of selective protein labeling inherent in the system, our results indicate that BONCAT has high specificity and resolution in proteomic profiling for minor subpopulations and proteins with low overall absolute abundance. We found multiple pathways and genes characteristic of K-state regulated differentially from the entire population, either significantly up- or down-regulated. Proteins that are uncharacterized or previously known for functions irrelevant of K-state are highly abundant in the subpopulation, providing new insight toward their alternative functions critical for K-state cells and future investigation directions of K-state study.
- Book Chapter
11
- 10.1016/b978-0-12-822312-3.00023-0
- Jan 1, 2021
- Molecular Docking for Computer-Aided Drug Design
Chapter 8 - Computational Modeling of Protein Three-Dimensional Structure: Methods and Resources
- Research Article
13
- 10.1186/1741-7007-11-44
- Apr 15, 2013
- BMC Biology
Segment assembly, structure alignment and iterative simulation in protein structure prediction
- Research Article
10
- 10.3724/sp.j.1123.2024.01011
- Jun 1, 2024
- Se pu = Chinese journal of chromatography
Given continuous improvements in industrial production and living standards, the analysis and detection of complex biological sample systems has become increasingly important. Common complex biological samples include blood, serum, saliva, and urine. At present, the main methods used to separate and recognize target analytes in complex biological systems are electrophoresis, spectroscopy, and chromatography. However, because biological samples consist of complex components, they suffer from the matrix effect, which seriously affects the accuracy, sensitivity, and reliability of the selected separation analysis technique. In addition to the matrix effect, the detection of trace components is challenging because the content of the analyte in the sample is usually very low. Moreover, reasonable strategies for sample enrichment and signal amplification for easy analysis are lacking. In response to the various issues described above, researchers have focused their attention on immuno-affinity technology with the aim of achieving efficient sample separation based on the specific recognition effect between antigens and antibodies. Following a long period of development, this technology is now widely used in fields such as disease diagnosis, bioimaging, food testing, and recombinant protein purification. Common immuno-affinity technologies include solid-phase extraction (SPE) magnetic beads, affinity chromatography columns, and enzyme linked immunosorbent assay (ELISA) kits. Immuno-affinity techniques can successfully reduce or eliminate the matrix effect; however, their applications are limited by a number of disadvantages, such as high costs, tedious fabrication procedures, harsh operating conditions, and ligand leakage. Thus, developing an effective and reliable method that can address the matrix effect remains a challenging endeavor. Similar to the interactions between antigens and antibodies as well as enzymes and substrates, biomimetic molecularly imprinted polymers (MIPs) exhibit high specificity and affinity. Furthermore, compared with many other biomacromolecules such as antigens and aptamers, MIPs demonstrate higher stability, lower cost, and easier fabrication strategies, all of which are advantageous to their application. Therefore, molecular imprinting technology (MIT) is frequently used in SPE, chromatographic separation, and many other fields. With the development of MIT, researchers have engineered different types of imprinting strategies that can specifically extract the target analyte in complex biological samples while simultaneously avoiding the matrix effect. Some traditional separation technologies based on MIP technology have also been studied in depth; the most common of these technologies include stationary phases used for chromatography and adsorbents for SPE. Analytical methods that combine MIT with highly sensitive detection technologies have received wide interest in fields such as disease diagnosis and bioimaging. In this review, we highlight the new MIP strategies developed in recent years, and describe the applications of MIT-based separation analysis methods in fields including chromatographic separation, SPE, diagnosis, bioimaging, and proteomics. The drawbacks of these techniques as well as their future development prospects are also discussed.
- Abstract
- 10.1063/4.0000885
- Sep 1, 2025
- Structural Dynamics
Cryogenic electron microscopy (cryo-EM) has become an essential method in structural biology for resolving large macromolecular assemblies. As cryo-EM continues to expand into more challenging targets, such as flexible assemblies and large macromolecular structures, the availability of medium to low-resolution maps (5–10 Å) has increased. However, interpreting these maps continues to be challenging due to the low quality of density features at limited map resolutions and inaccuracies in the predicted atomic models. While structure prediction methods such as AlphaFold (Abramson et al. 2024; Jumper et al. 2021) have led to significant improvements in model availability and quality, these models frequently suffer from domain misorientation, inaccurate flexible regions, or errors in complex formation. Consequently, traditional global fitting or rigid-body fitting approaches frequently struggle to achieve accurate model fitting, especially when large conformational changes or model errors exist in the predicted model.To address these challenges, we developed DMcloud, a new method that performs local structure fitting to improve model accuracy in cryo-EM maps at 5–10 Å resolution. DMcloud is designed to address cases where AlphaFold2 (AF2) models contain accurate local domains but incorrect global orientations. By converting both the AF2 model and cryo-EM map into point clouds, DMcloud performs iterative local alignments and denoising to refine model placement. This approach enables the correction of domain orientation errors and improves overall model-map agreement. DMcloud builds on our earlier method, DiffModeler (Wang et al. 2024), which performs global fitting of protein complexes using diffusion-based backbone tracing and AlphaFold-guided model assembly. While DiffModeler performs global structure modeling, DMcloud provides a complementary solution focused on fine-grained local fitting.We evaluated DMcloud on 71 intermediate- and 50 high-resolution cryo-EM maps using AF2 models from the AlphaFold Protein Structure Database. DMcloud outperformed traditional fitting methods, particularly in low-resolution cases and where AF2 models contained significant structural discrepancies. Figure 1 shows examples of modeling results by DMcloud for high and intermediate-resolution maps. Each row corresponds to a different target: a. EMD-20815, b. EMD-23192, c. EMD-21536, and d. EMD-12221 and shows (i) the EM density map, (ii) the reference PDB structure, (iii) a comparison between the DMcloud model and the reference PDB structure, and (iv) a comparison between the AlphaFold2 model and the reference PDB structure after structural alignment. As shown in Figure 1, DMcloud successfully corrected misoriented domains and reduced false model regions.DMcloud offers an accurate, automated framework for cryo-EM model fitting and is particularly effective in challenging cases where predicted models contain local inaccuracies, such as domain misorientations or flexible regions. By focusing on local structure alignment using point cloud representations and iterative refinement, DMcloud enables more precise model placement than conventional global or rigid-body fitting approaches. Its ability to selectively identify and adjust only the well-supported regions of given structure models makes it useful for large assemblies and partially resolved complexes. The method is freely available through our web server at https://em.kiharalab.org, alongside other AI-based tools for cryo-EM maps.
- Research Article
52
- 10.1080/19420862.2023.2175319
- Feb 12, 2023
- mAbs
Advances in structural biology and the exponential increase in the amount of high-quality experimental structural data available in the Protein Data Bank has motivated numerous studies to tackle the grand challenge of predicting protein structures. In 2020 AlphaFold2 revolutionized the field using a combination of artificial intelligence and the evolutionary information contained in multiple sequence alignments. Antibodies are one of the most important classes of biotherapeutic proteins. Accurate structure models are a prerequisite to advance biophysical property predictions and consequently antibody design. Specialized tools used to predict antibody structures based on different principles have profited from current advances in protein structure prediction based on artificial intelligence. Here, we emphasize the importance of reliable protein structure models and highlight the enormous advances in the field, but we also aim to increase awareness that protein structure models, and in particular antibody models, may suffer from structural inaccuracies, namely incorrect cis-amide bonds, wrong stereochemistry or clashes. We show that these inaccuracies affect biophysical property predictions such as surface hydrophobicity. Thus, we stress the importance of carefully reviewing protein structure models before investing further computing power and setting up experiments. To facilitate the assessment of model quality, we provide a tool “TopModel” to validate structure models.