ZINC 15--Ligand Discovery for Everyone.
Many questions about the biological activity and availability of small molecules remain inaccessible to investigators who could most benefit from their answers. To narrow the gap between chemoinformatics and biology, we have developed a suite of ligand annotation, purchasability, target, and biology association tools, incorporated into ZINC and meant for investigators who are not computer specialists. The new version contains over 120 million purchasable “drug-like” compounds – effectively all organic molecules that are for sale – a quarter of which are available for immediate delivery. ZINC connects purchasable compounds to high-value ones such as metabolites, drugs, natural products, and annotated compounds from the literature. Compounds may be accessed by the genes for which they are annotated as well as the major and minor target classes to which those genes belong. It offers new analysis tools that are easy for nonspecialists yet with few limitations for experts. ZINC retains its original 3D roots – all molecules are available in biologically relevant, ready-to-dock formats. ZINC is freely available at http://zinc15.docking.org.
- Book Chapter
3
- 10.1093/acrefore/9780199384655.013.606
- Jan 30, 2020
- Oxford Research Encyclopedia of Linguistics
The term “part of speech” is a traditional one that has been in use since grammars of Classical Greek (e.g., Dionysius Thrax) and Latin were compiled; for all practical purposes, it is synonymous with the term “word class.” The term refers to a system of word classes, whereby class membership depends on similar syntactic distribution and morphological similarity (as well as, in a limited fashion, on similarity in meaning—a point to which we shall return). By “morphological similarity,” reference is made to functional morphemes that are part of words belonging to the same word class. Some examples for both criteria follow: The fact that in English, nouns can be preceded by a determiner such as an article (e.g., a book, the apple) illustrates syntactic distribution. Morphological similarity among members of a given word class can be illustrated by the many adverbs in English that are derived by attaching the suffix –ly, that is, a functional morpheme, to an adjective (quick, quick-ly). A morphological test for nouns in English and many other languages is whether they can bear plural morphemes. Verbs can bear morphology for tense, aspect, and mood, as well as voice morphemes such as passive, causative, or reflexive, that is, morphemes that alter the argument structure of the verbal root. Adjectives typically co-occur with either bound or free morphemes that function as comparative and superlative markers. Syntactically, they modify nouns, while adverbs modify word classes that are not nouns—for example, verbs and adjectives. Most traditional and descriptive approaches to parts of speech draw a distinction between major and minor word classes. The four parts of speech just mentioned—nouns, verbs, adjectives, and adverbs—constitute the major word classes, while a number of others, for example, adpositions, pronouns, conjunctions, determiners, and interjections, make up the minor word classes. Under some approaches, pronouns are included in the class of nouns, as a subclass. While the minor classes are probably not universal, (most of) the major classes are. It is largely assumed that nouns, verbs, and probably also adjectives are universal parts of speech. Adverbs might not constitute a universal word class. There are technical terms that are equivalents to the terms of major versus minor word class, such as content versus function words, lexical versus functional categories, and open versus closed classes, respectively. However, these correspondences might not always be one-to-one. More recent approaches to word classes don’t recognize adverbs as belonging to the major classes; instead, adpositions are candidates for this status under some of these accounts, for example, as in Jackendoff (1977). Under some other theoretical accounts, such as Chomsky (1981) and Baker (2003), only the three word classes noun, verb, and adjective are major or lexical categories. All of the accounts just mentioned are based on binary distinctive features; however, the features used differ from each other. While Chomsky uses the two category features [N] and [V], Jackendoff uses the features [Subj] and [Obj], among others, focusing on the ability of nouns, verbs, adjectives, and adpositions to take (directly, without the help of other elements) subjects (thus characterizing verbs and nouns) or objects (thus characterizing verbs and adpositions). Baker (2003), too, uses the property of taking subjects, but attributes it only to verbs. In his approach, the distinctive feature of bearing a referential index characterizes nouns, and only those. Adjectives are characterized by the absence of both of these distinctive features. Another important issue addressed by theoretical studies on lexical categories is whether those categories are formed pre-syntactically, in a morphological component of the lexicon, or whether they are constructed in the syntax or post-syntactically. Jackendoff (1977) is an example of a lexicalist approach to lexical categories, while Marantz (1997), and Borer (2003, 2005a, 2005b, 2013) represent an account where the roots of words are category-neutral, and where their membership to a particular lexical category is determined by their local syntactic context. Baker (2003) offers an account that combines properties of both approaches: words are built in the syntax and not pre-syntactically; however, roots do have category features that are inherent to them. There are empirical phenomena, such as phrasal affixation, phrasal compounding, and suspended affixation, that strongly suggest that a post-syntactic morphological component should be allowed, whereby “syntax feeds morphology.”
- Conference Article
24
- 10.1109/iccv48922.2021.00016
- Oct 1, 2021
Neural networks trained with class-imbalanced data are known to perform poorly on minor classes of scarce training data. Several recent works attribute this to over-fitting to minor classes. In this paper, we provide a novel explanation of this issue. We found that a neural network tends to first under-fit the minor classes by classifying most of their data into the major classes in early training epochs. To correct these wrong predictions, the neural network then must focus on pushing features of minor class data across the decision boundaries between major and minor classes, leading to much larger gradients for features of minor classes. We argue that such an under-fitting phase over-emphasizes the competition between major and minor classes, hinders the neural network from learning the discriminative knowledge that can be generalized to test data, and eventually results in over-fitting. To address this issue, we propose a novel learning strategy to equalize the training progress across classes. We mix features of the major class data with those of other data in a mini-batch, intentionally weakening their features to prevent a neural network from fitting them first. We show that this strategy can largely balance the training accuracy and feature gradients across classes, effectively mitigating the under-fitting then over-fitting problem for minor class data. On several benchmark datasets, our approach achieves the state-of-the-art accuracy, especially for the challenging step-imbalanced cases.
- Supplementary Content
9
- 10.3389/fphar.2010.00005
- May 28, 2010
- Frontiers in Pharmacology
Natural products are of significant interest to us in many ways. The air we breathe, the water we drink, the foods we eat are all natural products. Many drug products and toxins are also “natural”. Natural product sciences now face challenges at many fronts. At a global and economical level, biodiversity is dimin-ishing everyday as the rain forest gives away to farmland and the coral reef is destroyed by pollution. As a result, many potentially valuable natural products are lost forever before we even know their very existence. An immediate implica-tion beyond the direct loss is that we have less natural products to “copy” from. We pharmacolo-gists may not have the power to change the world, but we can certainly contribute. At a social level, natural products are generally perceived as safe and largely devoid of side effects. This belief has been taken advantage of to promote “natural health products”. Beneficial effects of some are substantiated, but the only benefit of great many is only psychological. Some even created much damage to the believers. Aristolochic acid nephropathy is but only one example. Pharmacologists can, and are obliged to educate the public. At a scientific level, pharmacologists who are interested in natural products also face many challenges. First, rapid advances in related scientific disciplines (e.g. molecular biology, immunology) made us look like dino-saurs. We need to learn new concepts and technologies, and more importantly, to incorporate what we learn in our re-search. Second, modern chemistry and material sciences have blurred the boundary between natural products and synthetic materials. A great many of substances are neither natural nor purely man-made. We need to venture out of our comfort zones to embrace such great opportunities. Third, the use of natural products by human beings to combat diseases and promote health has never been with sin-gle known chemical identity until very recently. Many substances/recipes are known to have solid impact on human health; some have been validated by good clinical practice. Acceptance of application for herbal medications with multi-ple ingredients by the US FDA signals to the world a change of attitude towards acknowledging “good clinical observation”. This change probably also reflects a much more fundamental issue: the human body is a very complex organism. Very few diseases could be traced to a single cause. Multiple causes and pathological changes un-der most diseases call for intervention at multiple sites. One commonly encountered phenomenon in this area is the loss of biological activity in search of the “active compound”. Some, and probably most, are due to technical reasons, but we need to consider the alternative: 1 + 1 may be much more than 2. It is a general policy that studies accepted by this journal must be of known chemical entity or entities. This policy is needed to ensure the scientific quality but creates a set of problems. Frankly, we do not have a clear strategy to deal with this dilemma. Last but not least, we also face challenges at a practical level. First, many natural products with biological activity have poor solubility in water. As a result, research on these natural products is often difficult. Second, natural products are often limited in supply. Semi or total synthesis is a discipline that attempts to address this issue. Third, many natu-ral compounds with biological activity have complex chemical structures. Considering the almost infinite combination of chemical modifications, investigating the structure-activity relationship is tremendously time- and resource-consuming. Great challenges are often synonymous to great opportunities. The creation of this journal is a testimony to our collective effort in rising up to the challenges and turn opportunities to advances. We welcome submission of schol-arly manuscripts that may advance the pharmacology of natural sciences from every possible angle.
- Book Chapter
1
- 10.1201/9781003463429-6
- Mar 10, 2025
Natural products and their biological activities are currently a subject of great interest both for the scientific community and for the discipline of natural product chemistry, and the number of scientific studies in this field is increasing considerably. The plant kingdom is a rich reservoir of bioactive natural compounds. Other organisms, such as microorganisms, can also provide us with valuable biologically active products. The biological activity of natural products is concentrated in several species of plants and microorganisms. Many are dedicated to the search for natural compounds, in particular those that demonstrate biological, insecticidal, herbicidal, fungicidal, and biopesticidal activities that can be used as alternatives to the use of synthetic products. Some studies report the use of natural compounds against some infections alone or in combination. Natural products have been a source of inspiration for the production of new synthetic compounds both in the field of medicines and in the production of agrochemicals for crop protection. Due to the significant increase in the search for alternative products, essential oils can be used in the production of synthetic products for crop protection, as well as products synthesized from microorganisms that can also provide valuable active ingredients. Therefore, the aim of this research is to gather information on products from natural sources as well as new natural active ingredients for protection and safety in relation to synthetic products.
- Dissertation
- 10.53846/goediss-3730
- Jan 1, 2013
Stereochemistry is a very important subject in chemistry and life sciences, as most of the biologically active molecules, including proteins, oligonucleotides, sugars, lipids and small molecules exist in one chiral form. In addition, small molecule compounds with biological activity – drugs – with the same constitution yet different configuration can exhibit divergent biological activities. The determination of the relative configuration of stereocenters was first established by Fischer in 1891 via organic synthesis. The question of absolute configuration was not solved until the work of Bijvoet in 1951 with anomalous X-ray diffraction of single crystals. For the compounds which can neither be crystallized nor easily be converted into crystallizable derivatives, it can still be difficult to establish their stereochemistry. Within this thesis, the development and application of Residual Dipolar Coupling (RDC)-enhanced NMR techniques for the stereochemical elucidation of organic molecules – the RDC-based approach – will be presented. Residual dipolar couplings provide valuable information about the orientation of internuclear vectors with respect to the molecular frame, which complements nuclear Overhauser effects (NOEs), scalar J-couplings and chemical shifts that are short-range in nature. RDCs have been used for structure determination and dynamic studies of biomolecules for decades, however the use of this anisotropic parameter for configurational and conformational studies of small molecules has yet to be fully tapped. Herein, the new developments and improvements of this RDC-based approach for the study of small molecules can be categorized as: 1. Conformationally flexible molecules: Both single and multiple alignment tensor approaches were employed and compared. Molecular mechanics (MM) and molecular dynamics (MD) simulations were applied to statistically sample the extensive conformational space. 2. Natural products with limited availability: The preparation of a slim PH-gel for 1.7 mm NMR tubes has been developed for reproducibility and efficiency. RDC measurements in this low volume gel system have been made using a number of molecules. 3. Large numbers of unknown stereocenters: NOE and J-coupling restrained MD simulations with the floating chirality method were investigated. Combined with an RDC analysis in the final step, this enables the determination of the relative stereochemistry of a large number of stereocenters in a complex natural product. 4. Absolute configuration: An integrated approach is herein developed. This includes diastereomer differentiation using both isotropic and anisotropic NMR data. The chiroptical properties calculated using DFT of the NMR derivedensemble from this first step are used for enantiomer differentiation by comparison to experiments. Using these improved methods, the stereochemistry of six challenging naturalproducts – LLG1, naphth 1, naphth 2, comp.540a, vatiparol, and fibrosterol sulfateA – and one synthetic molecule – a Michael addition product – has been successfully established.
- Book Chapter
4
- 10.1007/11946465_10
- Jan 1, 2006
Many classifiers are designed with the assumption of well-balanced datasets. But in real problems, like protein classification and remote homology detection, when using binary classifiers like support vector machine (SVM) and kernel methods, we are facing imbalanced data in which we have a low number of protein sequences as positive data (minor class) compared with negative data (major class). A widely used solution to that issue in protein classification is using a different error cost or decision threshold for positive and negative data to control the sensitivity of the classifiers. Our experiments show that when the datasets are highly imbalanced, and especially with overlapped datasets, the efficiency and stability of that method decreases. This paper shows that a combination of the above method and our suggested oversampling method for protein sequences can increase the sensitivity and also stability of the classifier. Our method of oversampling involves creating synthetic protein sequences of the minor class, considering the distribution of that class and also of the major class, and it operates in data space instead of feature space. This method is very useful in remote homology detection, and we used real and artificial data with different distributions and overlappings of minor and major classes to measure the efficiency of our method. The method was evaluated by the area under the Receiver Operating Curve (ROC).KeywordsSupport Vector MachineReceiver Operating CharacteristicMinority ClassPositive InstanceImbalanced DataThese keywords were added by machine and not by the authors. This process is experimental and the keywords may be updated as the learning algorithm improves.
- Book Chapter
3
- 10.1007/978-3-540-71233-6_21
- Mar 12, 2007
Many classifiers are designed with the assumption of well-balanced datasets. But in real problems, like protein classification and remote homology detection, when using binary classifiers like support vector machine (SVM) and kernel methods, we are facing imbalanced data in which we have a low number of protein sequences as positive data (minor class) compared with negative data (major class). A widely used solution to that issue in protein classification is using a different error cost or decision threshold for positive and negative data to control the sensitivity of the classifiers. Our experiments show that when the datasets are highly imbalanced, and especially with overlapped datasets, the efficiency and stability of that method decreases. This paper shows that a combination of the above method and our suggested oversampling method for protein sequences can increase the sensitivity and also stability of the classifier. Synthetic Protein Sequence Oversampling (SPSO) method involves creating synthetic protein sequences of the minor class, considering the distribution of that class and also of the major class, and it operates in data space instead of feature space. We used G-protein-coupled receptors families as real data to classify them at subfamily and sub-subfamily levels (having low number of sequences) and could get better accuracy and Matthew's correlation coefficient than other previously published method. We also made artificial data with different distributions and overlappings of minor and major classes to measure the efficiency of our method. The method was evaluated by the area under the Receiver Operating Curve (ROC).
- Research Article
106
- 10.1021/jacs.6b12434
- Feb 15, 2017
- Journal of the American Chemical Society
The activation and selective transformation of virtually inexhaustible or easy-to-generate chemicals like N2, O2, CO2, CO, H2, or methane gas to value-added products is a lively area of current research, because of its economic relevance as well as its huge ecological impact. Biologists and chemists have put forth a lot of effort toward understanding and modeling the mechanisms of biological small-molecule activation, and in several catalytic cycles proposed for nickel-containing enzymes, nickel(I) plays a key role. In recent years also in synthetic chemistry the huge potential of complex nickel(I) units for the activation and transformation of small molecules has been discovered and exploited. This Perspective highlights some representative examples of nickel(I)-based small-molecule activation, intending to establish awareness of the competencies and scope of nickel(I) compounds.
- Research Article
261
- 10.1093/nar/gki456
- Jun 27, 2005
- Nucleic Acids Research
We present Babelomics, a complete suite of web tools for the functional analysis of groups of genes in high-throughput experiments, which includes the use of information on Gene Ontology terms, interpro motifs, KEGG pathways, Swiss-Prot keywords, analysis of predicted transcription factor binding sites, chromosomal positions and presence in tissues with determined histological characteristics, through five integrated modules: FatiGO (fast assignment and transference of information), FatiWise, transcription factor association test, GenomeGO and tissues mining tool, respectively. Additionally, another module, FatiScan, provides a new procedure that integrates biological information in combination with experimental results in order to find groups of genes with modest but coordinate significant differential behaviour. FatiScan is highly sensitive and is capable of finding significant asymmetries in the distribution of genes of common function across a list of ordered genes even if these asymmetries were not extreme. The strong multiple-testing nature of the contrasts made by the tools is taken into account. All the tools are integrated in the gene expression analysis package GEPAS. Babelomics is the natural evolution of our tool FatiGO (which analysed almost 22 000 experiments during the last year) to include more sources on information and new modes of using it. Babelomics can be found at .
- Research Article
1
- 10.1007/s12602-025-10796-9
- Nov 25, 2025
- Probiotics and antimicrobial proteins
The gut microbiome is a complex ecosystem of trillions of microbes residing in the human gastrointestinal tract. This microbial community plays a pivotal role in human health by swaying various physiological processes like metabolism, immunomodulation, antimicrobial protection, and psychological health. Natural and biological products comprising herbal extracts, dietary fibers, probiotics, prebiotics, and synbiotics have long been documented for their potential to modulate gut microbiota composition and function. Understanding the interactions between natural products and gut microbiota can provide valuable insights into their health-promoting effects and therapeutic potential. Literature was systematically sourced, analyzed, and compiled through searches across databases such as Science Direct, Google Scholar, and PubMed. The search terms include "gut microbiome," "Dysbiosis," "probiotics," "prebiotics," "synbiotics," "biological products," "gut health," "natural products," "dietary fibers," "resistant starches," "clinical study reports," "safety recommendations," etc., and several combinations of these words. This article provides a deep insight into the overview of natural and biological products aimed at promoting a healthy gut microbiome. It underscores the functional roles of the gut microbiota, including nutrient metabolism, antimicrobial defense, and gastrointestinal tract integrity. Furthermore, it delves into the detailed mechanisms of action and clinical evidence supporting the use of probiotics, prebiotics, synbiotics, herbal extracts, and dietary fibers in modulating gut microbiota composition and function. The biological activity, detailed mechanism of action, safety, and their considerable potential as modulators of gut microbiota composition and function have been extensively studied. Further research is necessary to clarify the underlying mechanisms of action to investigate their clinical implications.
- Conference Article
4
- 10.2118/171032-ms
- Oct 21, 2014
- SPE Eastern Regional Meeting
Horizontal drilling and hydraulic fracturing are commonplace in the shale plays across the United States with successful wells efficiently yielding large volumes of shale gas and oil at minimum time and cost. Typical well profiles involve building a curve and then drilling long horizontal sections to TD, with the horizontal section taking up to 7000ft to complete. One of the most cost effective means to drill such wells is with Positive Displacement Motor (PDM) Bottom Hole Assemblies (BHAs) consisting of PDMs with bend angles for directional control. When addressing drilling efficiency; a major contributor is BHA selection and the ability to select the most efficient combination of bit, PDM and stabilizer arrangement for each specific application. It is this selection process that forms the basis of this paper which details how a new suite of analysis tools were developed to analyse the dynamic behavior of a PDC bit on PDMs with variable bend angles and stabilizer configurations. Drilling on a bent motor in rotational mode has long proven challenging for both PDM and bit and the complex dynamic behavior has rarely been fully understood let alone analyzed. At the heart of the new suite of tools is the capability to analyse the complex trajectory of the cutting structure of the bit to simulate the Bottom Hole Pattern (BHP) created in the rock. Having understood the transient behavior of the bit, an analysis tool was then developed to analyse the drilling behavior of the bit and BHA which resulted in a new approach to PDC bit design for bent motor BHA applications with the design intent of providing smoother drilling, increased durability and higher rates of penetration. The first run of the new design of bit resulted in improved drilling performance, allowing the customer to drill the complete section to TD using a 2.38° bent motor BHA for the first time in the Marcellus. The negated need for multiple trips ultimately saved the customer up to 25% in completion time resulting in a cost saving of approximately $128k.
- Book Chapter
18
- 10.1007/978-3-319-73117-9_40
- Dec 22, 2017
Classification is one of the most fundamental and well-known tasks in data mining. Class imbalance is the most challenging issue encountered when performing classification, i.e. when the number of instances belonging to the class of interest (minor class) is much lower than that of other classes (major classes). The class imbalance problem has become more and more marked while applying machine learning algorithms to real-world applications such as medical diagnosis, text classification, fraud detection, etc. Standard classifiers may yield very good results regarding the majority classes. However, this kind of classifiers yields bad results regarding the minority classes since they assume a relatively balanced class distribution and equal misclassification costs. To overcome this problem, we propose, in this paper, a novel associative classification algorithm called Association Rule-based Classification for Imbalanced Datasets (ARCID). This algorithm aims to extract significant knowledge from imbalanced datasets by emphasizing on information extracted from minor classes without drastically impacting the predictive accuracy of the classifier. Experimentations, against five datasets obtained from the UCI repository, have been conducted with reference to four assessment measures. Results show that ARCID outperforms standard algorithms. Furthermore, it is very competitive to Fitcare which is a class imbalance insensitive algorithm.
- Research Article
90
- 10.1039/c8np00077h
- Jan 1, 2019
- Natural Product Reports
Covering: 1960s to end of August 2018 Plant glandular trichomes (GTs) are adaptive structures that are well known as "phytochemical factories" due to their impressive capacity to biosynthesize and store large quantities of specialized natural products. The natural products in GTs are chemically diverse and mostly function as defense chemicals, therefore GTs are frequently regarded as "the first defense line" of plants against biotic and abiotic stresses. More importantly, many GT natural products are commercially desirable, thanks to their significant biological activities, thus attracting extensive interest in their biosynthesis. Consequently, it is well known that plant GTs are not only important reservoirs of biologically active natural products but are also a valuable bank of novel biosynthetic genes and enzymes. The non-volatile or oxygenated natural products in plant GTs, which need longer biosynthetic pathways and more energy from the plants, are of particular interest due to their more extensive biological activities and high commercial value. This review mainly focuses on these non-volatile natural products in plant GTs, including their chemistry, biological activities and biosynthesis. The methods employed for investigating natural products and their biosynthesis in plant GTs are also comprehensively discussed.
- Research Article
52
- 10.1016/j.phytochem.2022.113532
- Dec 5, 2022
- Phytochemistry
Application of cinnamic acid in the structural modification of natural products: A review
- Research Article
53
- 10.1016/j.chembiol.2004.09.015
- Dec 1, 2004
- Chemistry & Biology
Macrolactamization of Glycosylated Peptide Thioesters by the Thioesterase Domain of Tyrocidine Synthetase