Articles published on Protein Data Bank
Authors
Select Authors
Journals
Select Journals
Duration
Select Duration
7932 Search results
Sort by Recency
- New
- Research Article
- 10.1042/bcj20260137
- Jul 8, 2026
- The Biochemical journal
- Joan Gizzio + 3 more
Humans have 437 catalytically competent protein kinase domains with a typical kinase fold resembling protein kinase A. The active form of a kinase must satisfy requirements for binding ATP, magnesium, and substrate. From structural bioinformatics analysis of 248 crystal structures of 54 kinase-substrate complexes, we derived structural criteria for the active form of typical protein kinases. We include well-known requirements on the DFG motif of the activation loop (ActLoop) and the N-terminal domain salt bridge but also on substrate-compatible states of the ActLoop N-terminal and C-terminal segments. With these criteria, only 123 of the 437 human catalytic protein kinases (cPKs) have active forms in the Protein Data Bank (PDB). Because the active forms are needed for understanding substrate specificity and mutational effects on catalytic activity in cancer and other diseases, we used AlphaFold2 to produce active models of all 437 human cPKs. This was accomplished with PDB templates that resemble substrate-bound structures, shallow sequence alignments of close paralogs/orthologs, and application of the active-kinase criteria to the output models. We selected models for each kinase based on intramolecular ActLoop ipSAE scores and showed that the highest-scoring models tend to have the lowest root mean square deviation (RMSD) to substrate-bound PDB structures. In a benchmark of 117 kinases, 92% have a highest-scoring AlphaFold2 model with backbone RMSD <2.0 Å to their benchmark active structure. Models for all 437 cPKs are available at https://dunbrack.fccc.edu/kincore/activemodels. We believe they may be useful for interpreting mutation-induced constitutive activity and as templates for modeling substrate and inhibitor binding to the active state.
- New
- Research Article
- 10.1107/s2059798326003943
- Jul 1, 2026
- Acta crystallographica. Section D, Structural biology
- Samuel P Foster + 4 more
Radiation damage to macromolecular structures remains a significant challenge for accurate structure solution by X-ray crystallography, leading to incorrect structural and chemical interpretation of the data. Site-specific radiation damage is insidious, typically unidentifiable solely from summary statistics, and is primarily discussed with reference to the predominant forms: disulfide-bond cleavage, metal-centre reduction and decarboxylation of acidic residues. A method is presented for identifying potentially oxidatively damaged cysteines by interrogating the accuracy of the built model within the electron density and the geometry of the difference density peaks surrounding a cysteine. We also highlight that cysteines located within protein active sites or that are in hydrolases are predisposed to this damage.
- New
- Research Article
- 10.1107/s2059798326005723
- Jul 1, 2026
- Acta crystallographica. Section D, Structural biology
- Airlie J Mccoy + 3 more
The application Scotty, implemented within the Phasertng codebase, was used to perform a PDB-wide analysis of lattice coincidences. Using a broad definition of lattice coincidence, the number of distinct lattice clusters is approximately half the total number of crystallographic PDB entries. In over one thousand lattice clusters entries are reported in different space groups, consistent with pseudo-symmetric variation within a common lattice framework. Space-group frequencies computed at the lattice-coincidence level update those obtained by entry-based counting, and more accurately reflect priors for novel crystal forms. Combining lattice clustering with sequence identity and deposited oligomeric annotations reveals multiple cases of inconsistent biological assembly assignments among structures sharing near-identical lattices, suggesting an opportunity to improve annotations. The survey also identifies a range of protein systems forming extended in cellulo paracrystalline arrays, including storage, sequestration, toxin and membrane-associated proteins, an understudied area of structural biology. Overall, the results demonstrate that lattice-level analysis provides a valuable perspective on macromolecular self-association.
- New
- Research Article
- 10.1038/s41598-026-59419-5
- Jun 30, 2026
- Scientific reports
- Muharib Alruwaili + 4 more
The emergence of multidrug-resistant (MDR) Salmonella enterica poses a critical threat to global health due to its capacity to evade conventional antibiotic therapies. In this study, we employed an integrated in silico strategy combining network pharmacology, molecular docking, ADMET profiling, and molecular simulations to explore diverse metabolites from gut microbiota and natural products as potential inhibitors. Antibiotic resistance genes were retrieved from the Comprehensive Antibiotic Resistance Database and screened via the Antibiotic Resistance Ontology, followed by essentiality analysis using the Database of Essential Genes. Among the identified targets, GyrA and AcrR emerged as key hub proteins through STRING-based protein-protein interaction analysis and pathway enrichment studies. The three-dimensional structures of AcrR and GyrA were obtained from the Protein Data Bank and Swiss-Model respectively,for structural evaluation. Metabolites were collected from the gutMGene and NPASS databases and subjected to docking analysis. Compounds exhibiting binding affinities below - 6.5kcal/mol were further evaluated for toxicity and pharmacokinetic properties. Penicillin G and indoxyl sulfate metabolite demonstrated favorable ADMET profiles and stable binding interactions, with molecular simulations supporting the structural stability of the ligand-protein complexes. These findings highlight the biodiversity of microbiota and plant-derived metabolites as promising resources against MDR S. enterica, warranting further experimental validation.
- New
- Research Article
- 10.1186/s12859-026-06448-6
- Jun 23, 2026
- BMC bioinformatics
- Jiyeon Min + 2 more
Iron-sulfur (Fe-S) clusters are ubiquitous cofactors in metalloproteins, supporting essential biological functions such as electron transfer, enzymatic catalysis, and metabolic regulation. Despite their importance, large-scale identification of Fe-S proteins remains challenging due to limitations of experimental methods and inconsistent annotations in existing databases. To address this, we introduced FeSseqdb, a curated sequence-level database derived from the Protein Data Bank (PDB), in which Fe-S cluster-containing chains are systematically verified using atomic coordinates. By standardizing diverse ligand annotations, FeSseqdb provides a unified and reliable resource for Fe-S protein research. Building on this foundation, a machine learning framework was developed to predict Fe-S proteins using only sequence-derived features, including amino acid composition, cysteine-related metrics, and sequence length. Among the classifiers evaluated, the random forest model trained on data resampled with SVMSMOTE achieved the highest predictive performance, highlighting the discriminative power of simple sequence features. To elucidate the biological relevance of these features, explainable AI methods were applied to identify key sequence characteristics associated with Fe-S proteins. Cysteine frequency and spatial distribution, along with proline content, emerged as primary contributors, which is consistent with their known structural roles in cluster coordination. Additionally, serine, glutamic acid, and arginine were identified as secondary determinants, in line with their reported roles in redox and electrostatic environments surrounding metal cofactors. The inclusion of these biologically relevant features demonstrates the potential of sequence-based models not only for accurate prediction but also for uncovering functional insights that align with known biochemical principles. This approach provides a foundation for large-scale, sequence-based discovery of Fe-S proteins and supports future investigations into their functional diversity across proteomes.
- New
- Research Article
- 10.1021/acs.jcim.6c00846
- Jun 22, 2026
- Journal of chemical information and modeling
- M Andrés Velasco-Saavedra + 10 more
Covalent binders (CBs) represent a broad and increasingly important class of compounds with substantial clinical potential. Although they have historically been overlooked because of toxicity concerns, covalent inhibitors are now recognized as a valuable therapeutic strategy with important pharmacological implications. While the chemical diversity of CBs has been well documented, their comprehensive characterization in terms of chemical space and enzyme class remains underexplored. In this study, we present a systematic analysis of CBs available in the Protein Data Bank (PDB) grouped by their biological target. Our extensively curated database includes 3,585 chemical structures deposited in the PDB, significantly expanding the chemical space representation of CBs. We explore this space using therapeutically relevant targets in medicinal chemistry as case studies, highlighting the breadth and specificity of covalent interactions captured in the updated dataset. Furthermore, we compare the chemical space of CBs with key reference datasets, including approved drugs, natural products, and fragment libraries. This comparison underscores the unique and complementary roles that CBs play in drug discovery. The updated database of CBs, freely available to the scientific community, provides a valuable resource to support the rational design and development of covalently binding therapeutic agents.
- New
- Research Article
- 10.1021/acs.jcim.5c02404
- Jun 22, 2026
- Journal of chemical information and modeling
- Suparno Nandi + 1 more
New and improved methods for visualizing complex macromolecules in atomic detail continue to expand structural information in the Protein Data Bank but accurately refining atomic models from experimental maps remains a challenge due to efficiency limitations of current refinement approaches. Standard PHENIX refinement can partially address these limitations with its speed and accessibility but often fails to yield the best model compared to more computationally demanding approaches. To support improved macromolecular model building, we therefore developed "KNexPHENIX", a PHENIX-based workflow that combines staged refinement, geometry minimization, and customized refinement parameters. KNexPHENIX can be used to refine macromolecular structures obtained via cryo-electron microscopy (cryo-EM) or X-ray crystallography, regardless of molecular size or composition. KNexPHENIX was evaluated on deposited structures and de novo models and consistently produced models with lower MolProbity scores, indicating improved model stereochemistry, compared to default PHENIX, REFMAC Servalcat, REFMAC, or CERES refinement. Importantly, this was accomplished while maintaining model-to-map correlation for cryo-EM data sets and maintaining or reducing the Rfree - Rwork difference below accepted thresholds for X-ray crystallographic structures, thus limiting overfitting while preserving refinement accuracy. While remaining dependent on the initial model and the choice of starting structure (e.g., from AlphaFold, Boltz, or RoseTTAFold), these results establish the KNexPHENIX workflow as a practical, accessible approach for refining both cryo-EM and crystallographic structures, enabling the generation of improved models for deposition.
- New
- Research Article
- 10.64898/2026.06.20.733508
- Jun 20, 2026
- bioRxiv : the preprint server for biology
- Jad Zbib + 1 more
In the evolving landscape of RNA research, the classification and analysis of RNA motifs is necessary to uncover the intricate mechanisms governing cellular and viral processes. Here we apply the coarse-grained RAG (RNA-As-Graphs) framework to advance the classification and understanding of RNA motifs, with a focus on expanding the RNA Motif Atlas through the inclusion of novel viral RNA structures. By analyzing 273 experimentally resolved viral RNA structures from the Protein Data Bank (PDB) using RAG dual-graph representations, we identify 14 previously uncatalogued viral RNA motifs. These motifs, which include tRNA mimicking domains, exoribonuclease-resistant domains, and internal ribosome entry sites, expand the diversity of RNA to a total number of 197 dual graph motifs. We applied k-means, PAM, and Ward clustering and observed substantial overlap between viral and general RNAs. The expanded library of RNA motifs and submotifs provides a resource for motif discovery and RNA design.
- New
- Research Article
- 10.54503/0321-1339-2026.126.2-2
- Jun 19, 2026
- Reports NAS RA
- Hamlet Khachatryan
Human Immunodeficiency Virus-1 protease (HIV-1 PR) is among the most extensively studied drug targets in the Protein Data Bank (PDB), with more than 600 structural models predominantly derived by X-ray crystallography. This study presents a comprehensive analysis of the binding-site conformational space across the available structural record: 690 crystal structures deposited in the PDB with ≥90% sequence identity and resolution better than or equal to 2.50 Å, of which 684 were successfully featurized by pipeline. The structural dataset covers wild-type enzyme, crystallographic stabilization mutants, drug-resistant variants, and 452 distinct inhibitor binders. Each binding site was featurized as a volume-filling point cloud with six descriptors (electrostatic potential, lipophilicity, and pharmacophoric features) and represented as a geodesic distance matrix with further embedding in spectral distance space. Affinity propagation clustered all pockets into 16 discrete conformational states, with four dominant states accounting for 85% of all structures.
- New
- Research Article
- 10.1021/acs.biochem.6c00170
- Jun 18, 2026
- Biochemistry
- Nandita Puri + 1 more
Lipid-protein interactions are ubiquitous in biology, where they are fundamental to membrane structure, cell signaling, immunology, and metabolism. Despite the availability of thousands of experimentally determined lipid-protein structures, the molecular basis for lipid recognition and specificity across the lipid-protein interactome remains incompletely understood. Here, we report a systematic analysis of 113,782 annular and nonannular lipid-protein complexes spanning the eight lipid classes. Pairwise atomic interactions are linked to lipid and protein physicochemical properties and binding geometries. Hydrophobic contacts, hydrogen bonds, and salt bridges contributed to over 99% of lipid-protein interactions. Lipid class-, protein sublocalization-, protein function-, and protein fold-dependent trends were identified. Protein pockets were finely tuned for lipid size, shape, and polarity: fatty acyls associated with narrow, moderately hydrophobic pockets; saccharolipids and glycerophospholipids bound to larger, polar cavities; and sterols and prenols preferentially occupied compact hydrophobic sites. Global analysis across different protein families identified similarities in interaction profiles, while also highlighting protein-specific recognition adapted to biochemical function. Lipid-protein interaction maps were projected onto lipid structures to uncover conserved and divergent hotspots and coldspots across lipid classes. The heatmaps imply that recognition and specificity are mediated by tailored anchoring of polar head groups and varying interaction with hydrophobic tails. Together, the data establish nature's principles governing lipid binding, lipid selectivity, and complex stability, and collectively provide a molecular atlas of the lipid-protein interactome. The work enables the elucidation of lipid biology at scale and establishes guiding principles for the rational design of chemical probes and therapeutics targeting lipid biology.
- New
- Research Article
- 10.1021/acs.jpcb.6c01344
- Jun 18, 2026
- The journal of physical chemistry. B
- Vishnu Santosh Kumar + 3 more
Disulfide bonds play essential structural and regulatory roles in proteins; however, the noncovalent interactions (NCIs) between these disulfide motifs remain largely unexplored. In this work, we investigate the nature of S-S···S-S interactions using a combined Protein Data Bank (PDB) survey and quantum chemical calculations. Statistical analyses of protein crystal structures reveal the preferred intermolecular sulfur-sulfur distances of approximately 3.6 Å and 4.8 Å, suggesting the presence of stabilizing NCIs between disulfide bridges. To elucidate the intrinsic features of these interactions, dimethyl disulfide (DMDS) was employed as a model system. Computational analysis of multiple dimer conformers reveals that for arrangements with intermolecular sulfur-sulfur distance greater than 3.6 Å (sum of van der Waals radii of two sulfurs), stability is primarily driven by C-H···S hydrogen bonds. In contrast, for conformers with shorter intermolecular sulfur-sulfur distances (⩽3.6 Å), stabilization arises from a cooperative interplay between C-H···S hydrogen bonds and directional S···S chalcogen bonds. This study presents the first report of S···S chalcogen bonding between disulfide bonds, highlighting its potential role in biochemical architecture.
- New
- Research Article
- 10.1371/journal.pone.0351662
- Jun 16, 2026
- PLOS One
- Sevilay Gülesen + 3 more
Viruses such as coronaviruses or filoviruses use their surface glycoproteins (GPs) to attach to the host cell, triggering the fusion of the viral membrane with the endosome membrane. Epitopes on the viral GP are major targets for antibody-mediated recognition and neutralization. During the fusion process, the GP undergoes conformational changes triggered by fluctuations in environmental pH. Structural states are typically classified into three distinct conformations: prefusion, intermediate, and postfusion. These conformations serve as essential templates for prediction of conformational epitopes and structure-based vaccine design. Despite their importance, many viral GP structures remain absent from the Protein Data Bank (PDB). Fortunately, recent breakthroughs in computational structure prediction have greatly enhanced the accuracy and accessibility of protein modeling. In this study, we utilized AlphaFold2-Multimer (AF2-M), version 2.3, to predict various GP structural conformations and observed that the overall frequency of predictions in the postfusion conformation is low. Therefore, we hypothesized that adapting the AF2-M protocol is necessary to enrich for specific conformations, thereby enabling the prediction of both pre- and postfusion conformations. AF2-M requires only the input sequence and internally generates multiple sequence alignments (MSAs) and optional templates before applying its pretrained model weights. We tested the use of template data to enrich pre- or postfusion conformations and demonstrated that our approach significantly increases the prediction frequency of class I fusion protein structures in both conformations, with the template dataset playing a crucial role in guiding modeling towards the intended state. Furthermore, we showed that the lack of correlation between pLDDT and TM-scores suggests that low pLDDT values may obscure the presence of valid alternative conformations.
- New
- Research Article
- 10.1021/acs.jcim.6c00352
- Jun 15, 2026
- Journal of chemical information and modeling
- Edward H Snell + 11 more
Accurate representation of metal ions in macromolecular structures is critical for chemical interpretation, computational modeling, and machine-learning methods that rely on Protein Data Bank (PDB) entries. However, the elemental identity of metals modeled in crystallographic structures is often inferred indirectly and rarely validated experimentally. Here, we combine Particle Induced X-ray Emission (PIXE) and X-ray Fluorescence Spectroscopy (XRFS) to determine the elemental composition of protein samples used to generate 70 deposited metalloprotein crystal structures. By analyzing the original protein material employed for crystallization, but before the addition of crystallization buffer solutions, we assess whether the modeled metal ions in deposited structures are consistent with experimentally detectable elemental content. We find that in a majority of cases, the metals modeled in the corresponding PDB entries are inconsistent with the metals present in the protein samples before crystallization, or that additional metals are present but not represented in the structural models. Spectroscopic results were integrated with automated crystallographic validation metrics, including real-space Z-difference (RSZD) analysis and systematic rerefinement, to evaluate atomic-number mismatch at metal sites. PIXE and XRFS show strong agreement for dominant elemental signals and provide complementary, scalable approaches for identifying suspect metal assignments. This work does not address physiological or functional metalation but instead highlights a widespread data integrity issue in deposited macromolecular structures, PDB-wide. These results establish an experimentally corroborated link between elemental identity and crystallographic validation metrics, enabling the large-scale detection of chemically inconsistent annotations in structural databases used for computational modeling and machine learning.
- Research Article
- 10.1080/09553002.2026.2686688
- Jun 12, 2026
- International Journal of Radiation Biology
- S Ronan Fisher + 3 more
Purpose Human calmodulin-1 (CaM) undergoes calcium-dependent conformational changes that are difficult to capture owing to substantial structural flexibility. This study aims to show that double electron-electron resonance (DEER) spectroscopy can measure ensembles of CaM conformations. The most probable distances are compared to structural models from experimental structures and deep-learning predictions. Materials and Methods Four double-cysteine variants of a recombinant CaM were engineered for site-directed spin labeling. Spin-labeled CaM was measured with DEER spectroscopy in both the presence (holo CaM) and absence (apo CaM) of Ca(II). Distance distributions were compared with predictions derived from Protein Data Bank structures and models generated from AlphaFold2 and IntFOLD7. DEERefiner was used to generate conformers from these models using the distance-distributions as restraints. Results The distance distributions from DEER were consistent with conformational heterogeneity. Distance distributions differed between apo and holo conditions. The results with Ca(II) were consistent with an ensemble of conformations including the canonical crystallographic structure (1CLL). Apo state distance distributions were not reproduced well by the models. Upon refinement with the distance distributions the models aligned better with the most probable distances from the apo and holo results. Comparisons of the refined models indicated that Ca(II) binding causes subtle rearrangements between EF-hand domains. Conclusions The DEER-derived distance distributions for the four doubly-labeled CaM variants demonstrate that apo and holo CaM populate broad but different conformational ensembles in frozen solution. Comparison of apo and holo distributions indicate that Ca(II) binding induces a subtle, yet significant, structural rearrangement. These results illustrate how DEER-guided modeling can provide deeper insight into flexible protein ensembles that are not captured by existing static structures.
- Research Article
- 10.1016/j.envres.2026.125000
- Jun 11, 2026
- Environmental research
- Nazmi Harith-Fadzilah + 3 more
In Silico Identification and characterisation of putative biphenyl degradation mechanism in gut-Derived Pediococcus pentosaceus.
- Research Article
- 10.1073/pnas.2607035123
- Jun 9, 2026
- Proceedings of the National Academy of Sciences
- Jing Yang + 2 more
Generative AI algorithms such as the transformer and diffusion models have greatly empowered de novo design of proteins capable of specifically interacting with designated structural sites on another protein. Most of these design methods employ a top-down approach, in which an overall protein shape is generated by an AI model to pack against a given structural site, followed by sequence design to optimize the interaction. Despite being trained on limited protein complex structures available in the database, the top-down approach has yielded encouraging results. Here, we propose a bottom-up approach that generates atom clusters for optimal packing against a specified structured region for informing the design of protein-protein interactions. To this end, we trained a masked discrete diffusion model, named Void-X, that uses the diffusion transformer to learn atomic-level interactions and fill atomic voids in protein interaction interfaces. Void-X was trained using 8.7 million spherical clusters of atoms from experimental structures in the Protein Data Bank. In each cluster, ~70% of the atoms are used as context (or prompt), and ~30% are masked for information recovery (or answer). By training the model with 172 million parameters, Void-X achieves an overall accuracy of 78.3% and 68.2% for intra- and interchain spherical clusters, respectively. Furthermore, we find that information entropy is a reliable indicator of the prediction accuracy for Void-X. This level of performance allows de novo generation of molecular interactions at the atomic level, offering an alternative approach of protein design complementary to the existing ones.
- Research Article
- 10.2174/0115701638439592260516051944
- Jun 8, 2026
- Current drug discovery technologies
- Sujeet Maurya + 3 more
The landscape of drug discovery is being rapidly transformed by the integration of computational intelligence (CI) techniques with big data resources in medicinal chemistry. Traditional drug development methods are often time-consuming, costly, and prone to high attrition rates. In contrast, data-driven and algorithmic approaches-powered by machine learning, deep learning, and hybrid models-enable rapid, precise, and predictive decision-making across all phases of drug design. This review provides a comprehensive overview of how CI techniques, including supervised and unsupervised learning, convolutional and recurrent neural networks, evolutionary algorithms, and fuzzy logic systems, are reshaping drug discovery pipelines. We explore the role of massive chemical and biological databases such as PubChem, ChEMBL, DrugBank, and the Protein Data Bank, while also highlighting the importance of data quality, curation, and standardization. The integration of computational tools in drug discovery is discussed across key stages-target identification, hit discovery, lead optimization, and de novo molecular design-supported by examples such as AlphaFold, Atomwise, and In silico Medicine. This review aims to offer a holistic understanding of how computational intelligence, when combined with robust data infrastructure, can significantly accelerate the discovery and development of safer, more effective drugs.
- Research Article
- 10.1016/j.fsi.2026.111484
- Jun 2, 2026
- Fish & shellfish immunology
- Xueting Zhao + 6 more
A structurally conserved Attacin mediates IMD-dependent intestinal immunity against Vibrio parahaemolyticus in Eriocheir sinensis.
- Research Article
- 10.1002/pro.70648
- Jun 2, 2026
- Protein Science : A Publication of the Protein Society
- Hana Pokojn\Xe1 + 4 more
Through the integration of science, art, and technology, we present Multimedia Resolution Data Artistic Educational Tool (MRDAET), an interactive installation designed to support the visualization and understanding of molecular reactions. The installation focuses on ATP synthesis and the electron transport chain—core biochemical processes that are often difficult to grasp due to their complexity and abstract nature. Inspired by David Goodsell's fusion of scientific accuracy and artistic expression, MRDAET employs physical 3D models derived from structural data in the Protein Data Bank, augmented with projection mapping and enhanced with illustrations. The installation encourages users to explore molecular processes through layered interactivity that combines tangible models with animated visual overlays. MRDAET was evaluated at an art and technology event, where user feedback indicated that the installation was both enjoyable and perceived as helpful in education. Participants reported improved understanding of the presented biochemical concepts and expressed interest in increased interactivity, particularly through Mixed Reality integration connecting physical models and dynamic animations. While individual and combined visual modalities have been shown to be effective in science education, immersive interactive installations of this type remain underexplored in this context. This paper presents MRDAET as the third iteration following prior designs, offering a science‐driven educational approach that warrants further study and application to additional scientific areas.
- Research Article
- 10.1016/j.sbi.2026.103243
- Jun 1, 2026
- Current opinion in structural biology
- Carlos A H Fernandes + 1 more
Membrane proteins (MPs) play essential roles in a wide range of cellular processes and represent major therapeutic targets. Nevertheless, their structural and functional characterization remains challenging due to inherent difficulties in production, extraction, and stabilization outside their native lipid environment. The rise of cryogenic electron microscopy (cryo-EM) has markedly accelerated the structure determination of MPs through single-particle analysis (SPA). A systematic examination of high-resolution cryo-EM SPA structures deposited in the Protein Data Bank (PDB) over the past two years provides a comprehensive overview of the most frequently used amphipathic environments. We discuss the strengths and limitations of each approach and underscore the ongoing need to develop near-native environment strategies to improve the interpretation of MP structures under membrane-like conditions.