Accelerate Literature Icon
Want to do a literature review? Try our new Literature Review workflow

Deep learning for assay nuisance compound detection using a gated co-attention graph embedding model (CAGE-Fusion).

  • Abstract
  • Literature Map
  • Similar Papers
Abstract
Translate article icon Translate Article Star icon

In drug discovery, nuisance compounds are chemicals that interfere with biochemical or cell-based assays, often generating misleading signals unrelated to true activity. These molecules operate through diverse, context-dependent mechanisms, ranging from colloidal aggregation to covalent reactivity. They are often referred to as Pan-Assay Interference Compounds (PAINS), signal interferers, or frequent hitters, depending on their mode of action, and they lead to costly false positives and resource misallocation. They are difficult to distinguish from genuine pharmacological activity using rigid substructure filters. Although deep learning has emerged as a powerful alternative, most existing architectures process molecular graphs or SMILES sequences in isolation. We introduce CAGE-Fusion (Co-Attention Graph Embedding Fusion), a multimodal framework that employs a gated co-attention mechanism to enforce bidirectional information exchange between graph-based and sequence-based encoders. This architecture enables the iterative refinement of atom and token embeddings, yielding a chemically coherent representation that captures cross-modal dependencies. CAGE-Fusion demonstrates competitive predictive performance across several MoleculeNet benchmarks and achieves a macro-averaged PR-AUC of 0.73 and ROC-AUC of 0.94 on a comprehensive nuisance compound classification task. Beyond prediction, the model visualizes attention-weighted regions that contribute to a classification decision.Scientific contributionWe present CAGE-Fusion, a gated co-attention architecture that enables bidirectional information exchange between graph-based encoders and SMILES representations to learn cross-modal molecular embeddings. We evaluate the framework on standardized molecular property benchmarks and apply it in a focused case study on assay nuisance compound detection across four nuisance categories: aggregation, luciferase inhibition, reactivity, and promiscuity.

Similar Papers
  • Research Article
  • Cite Count Icon 102
  • 10.1021/acs.jcim.8b00677
Hit Dexter 2.0: Machine-Learning Models for the Prediction of Frequent Hitters.
  • Jan 9, 2019
  • Journal of Chemical Information and Modeling
  • Conrad Stork + 3 more

Assay interference caused by small molecules continues to pose a significant challenge for early drug discovery. A number of rule-based and similarity-based approaches have been derived that allow the flagging of potentially "badly behaving compounds", "bad actors", or "nuisance compounds". These compounds are typically aggregators, reactive compounds, and/or pan-assay interference compounds (PAINS), and many of them are frequent hitters. Hit Dexter is a recently introduced machine learning approach that predicts frequent hitters independent of the underlying physicochemical mechanisms (including also the binding of compounds based on "privileged scaffolds" to multiple binding sites). Here we report on the development of a second generation of machine learning models which now covers both primary screening assays and confirmatory dose-response assays. Protein sequence clustering was newly introduced to minimize the overrepresentation of structurally and functionally related proteins. The models correctly classified compounds of large independent test sets as (highly) promiscuous or nonpromiscuous with Matthews correlation coefficient (MCC) values of up to 0.64 and area under the receiver operating characteristic curve (AUC) values of up to 0.96. The models were also utilized to characterize sets of compounds with specific biological and physicochemical properties, such as dark chemical matter, aggregators, compounds from a high-throughput screening library, drug-like compounds, approved drugs, potential PAINS, and natural products. Among the most interesting outcomes is that the new Hit Dexter models predict the presence of large fractions of (highly) promiscuous compounds among approved drugs. Importantly, predictions of the individual Hit Dexter models are generally in good agreement and consistent with those of Badapple, an established statistical model for the prediction of frequent hitters. The new Hit Dexter 2.0 web service, available at http://hitdexter2.zbh.uni-hamburg.de , not only provides user-friendly access to all machine learning models presented in this work but also to similarity-based methods for the prediction of aggregators and dark chemical matter as well as a comprehensive collection of available rule sets for flagging frequent hitters and compounds including undesired substructures.

  • Research Article
  • Cite Count Icon 33
  • 10.1080/17460441.2019.1654453
Dealing with frequent hitters in drug discovery: a multidisciplinary view on the issue of filtering compounds on biological screenings
  • Aug 16, 2019
  • Expert Opinion on Drug Discovery
  • Rafael Ferreira Dantas + 6 more

Introduction: The timely identification biologically active chemicals, in disease relevant screening assays, is a major endeavor in drug discovery. The existence of frequent hitters (FHs) in non-related assays poses a formidable challenge in terms of whether to consider these molecules as chemical gold or promiscuous non-selective reactive trash (also known as PAINS – pan assay interference compounds). Areas covered: In this review, the authors bring together expertize in synthetic chemistry, cheminformatics and biochemistry, three key areas for dealing with FHs. They discuss synthetic methods facilitating preparation of chemically diverse molecular libraries, while favoring activity in the biological space. They also survey and discuss recent computational advances in the prediction of PAINS from chemical structures. Finally, they review experimental approaches for the validation of the biological activity of screening hits and discuss alternatives for exploiting promiscuity and chemical reactivity. Expert opinion: It’s essential to develop more efficient computational methods to reliably recognize PAINS in distinct molecular environments. Accordingly, advances in synthetic chemistry hold the promise to provide a better quality of chemical matter for drug discovery. Medicinal chemists should be more open to screening for hits showing biologically complex mechanisms of action rather than discarding molecules that may prove valuable as innovative disease treatments.

  • Research Article
  • Cite Count Icon 1
  • 10.2174/1568026621666210512020434
Artificial Intelligence and Cheminformatics-Guided Modern Privileged Scaffold Research.
  • May 11, 2021
  • Current topics in medicinal chemistry
  • Han-Yue Qiu + 3 more

With the rapid development of computer science in scopes of theory, software, and hardware, artificial intelligence (mainly in form of machine learning and more complex deep learning) combined with advanced cheminformatics is playing an increasingly important role in drug discovery process. This development has also facilitated privileged scaffold-related research. By definition, a privileged scaffold is a structure that frequently occurs in diverse bioactive molecules, either has a diverse family affinity or is selective to multiple family members in a superfamily, whilst it is different from the"frequent hitters", or the "pan-assay interference compounds". The long history of the use of this concept has witnessed a functional shift from stand-alone technology towards an integrated component in the drug discovery toolbox. Meanwhile, continuous efforts have been dedicated to deepening the understandings of the features of known privileged scaffolds. In this contribution, we focus on the current privileged scaffold-related research driven by state-of-art artificial intelligence approaches and cheminformatics. Representative cases with an emphasis on distinct research aspects are presented, including an update of the knowledge on privileged scaffolds, proofof- concept tools, and workflows to identify privileged scaffolds and to carry on de novo design, informatic SAR models with diversely complex data sets to provide an instructive prediction on new potential molecules bearing privileged scaffolds.

  • Research Article
  • Cite Count Icon 14
  • 10.1021/acs.jcim.8b00385
Exploring Activity Profiles of PAINS and Their Structural Context in Target-Ligand Complexes.
  • Aug 14, 2018
  • Journal of Chemical Information and Modeling
  • Vishal B Siramshetty + 2 more

Assay interference is an acknowledged problem in high-throughput screening, and pan-assay interference compounds (PAINS) filters are one of a number of approaches that have been suggested for identification of potential screening artifacts or frequent hitters. Many studies have highlighted that the unwary usage of these structural alerts should be reconsidered and criticized their extrapolation beyond the applicability domain. A large-scale investigation of the activity profiles and the structural context of PAINS might provide a better assessment of whether this extrapolation is valid. To this end, multiple publicly accessible compound collections were screened, and the PAINS statistics are comprehensively presented and discussed. Next, the promiscuity trends and activity profiles of PAINS were compared with those compounds not matching any PAINS substructures. Overall, PAINS demonstrated higher promiscuity and relatively higher assay hit rates compared with the other compounds. Furthermore, nearly 2000 distinct target-ligand complexes containing PAINS were analyzed, and the interactions were quantified and compared. In more than 50% of the instances, the PAINS atoms participated in interactions more frequently compared with the remaining atoms of the ligand structure. Many PAINS participated in crucial interactions that were often responsible for binding of the ligand, which reaffirms their distinction from those responsible for assay interference. In conclusion, we reinforce that while it is important to employ compound filters to eliminate nonspecific hits, establishing a set of statistically significant and validated PAINS filters is essential to restrain the black-box practice of triaging screening hits matching any of the proposed 480 alerts.

  • Research Article
  • Cite Count Icon 13
  • 10.1016/j.ailsci.2021.100007
Computational prediction of frequent hitters in target-based and cell-based assays
  • Aug 8, 2021
  • Artificial Intelligence in the Life Sciences
  • Conrad Stork + 2 more

Compounds interfering with high-throughput screening (HTS) assay technologies (also known as “badly behaving compounds”, “bad actors”, “nuisance compounds” or “PAINS”) pose a major challenge to early-stage drug discovery. Many of these problematic compounds are “frequent hitters”, and we have recently published a set of machine learning models (“Hit Dexter 2.0”) for flagging such compounds.Here we present a new generation of machine learning models which are derived from a large, manually curated and annotated data set. For the first time, these models cover, in addition to target-based assays, also cell-based assays. Our experiments show that cell-based assays behave indeed differently from target-based assays, with respect to hit rates and frequent hitters, and that dedicated models are required to produce meaningful predictions. In addition to these extensions and refinements, we explored a variety of additional setups for modeling, including the combination of four machine learning classifiers (i.e. k-nearest neighbors (KNN), extra trees, random forest and multilayer perceptron) with four sets of descriptors (Morgan2 fingerprints, Morgan3 fingerprints, MACCS keys and 2D physicochemical property descriptors).Testing on holdout data as well as data sets of “dark chemical matter” (i.e. compounds that have been extensively tested in biological assays but have never shown activity) and known bad actors show that the multilayer perceptron classifiers in combination with Morgan2 fingerprints outperform other setups in most cases. The best multilayer perceptron classifiers obtained Matthews correlation coefficients of up to 0.648 on holdout data. These models are available via a free web service.

  • Research Article
  • Cite Count Icon 42
  • 10.1021/acs.jcim.8b00104
Modeling Small-Molecule Reactivity Identifies Promiscuous Bioactive Compounds.
  • Jul 10, 2018
  • Journal of Chemical Information and Modeling
  • Matthew K Matlock + 3 more

Scientists rely on high-throughput screening tools to identify promising small-molecule compounds for the development of biochemical probes and drugs. This study focuses on the identification of promiscuous bioactive compounds, which are compounds that appear active in many high-throughput screening experiments against diverse targets but are often false-positives which may not be easily developed into successful probes. These compounds can exhibit bioactivity due to nonspecific, intractable mechanisms of action and/or by interference with specific assay technology readouts. Such "frequent hitters" are now commonly identified using substructure filters, including pan assay interference compounds (PAINS). Herein, we show that mechanistic modeling of small-molecule reactivity using deep learning can improve upon PAINS filters when modeling promiscuous bioactivity in PubChem assays. Without training on high-throughput screening data, a deep learning model of small-molecule reactivity achieves a sensitivity and specificity of 18.5% and 95.5%, respectively, in identifying promiscuous bioactive compounds. This performance is similar to PAINS filters, which achieve a sensitivity of 20.3% at the same specificity. Importantly, such reactivity modeling is complementary to PAINS filters. When PAINS filters and reactivity models are combined, the resulting model outperforms either method alone, achieving a sensitivity of 24% at the same specificity. However, as a probabilistic model, the sensitivity and specificity of the deep learning model can be tuned by adjusting the threshold. Moreover, for a subset of PAINS filters, this reactivity model can help discriminate between promiscuous and nonpromiscuous bioactive compounds even among compounds matching those filters. Critically, the reactivity model provides mechanistic hypotheses for assay interference by predicting the precise atoms involved in compound reactivity. Overall, our analysis suggests that deep learning approaches to modeling promiscuous compound bioactivity may provide a complementary approach to current methods for identifying promiscuous compounds.

  • Research Article
  • Cite Count Icon 12
  • 10.1002/cmdc.202100710
Promoting GAINs (Give Attention to Limitations in Assays) over PAINs Alerts: no PAINS, more GAINs.
  • Feb 11, 2022
  • ChemMedChem
  • Malcolm Z Y Choo + 1 more

Promoting GAINs (Give Attention to Limitations in Assays) over PAINs Alerts: no PAINS, more GAINs.

  • Research Article
  • Cite Count Icon 173
  • 10.1517/17460441.2012.688743
Rhodanine as a scaffold in drug discovery: a critical review of its biological activities and mechanisms of target modulation
  • May 19, 2012
  • Expert Opinion on Drug Discovery
  • Tihomir Tomašić + 1 more

Introduction: Rhodanine-based compounds have been associated with numerous biological activities. After many years of research in drug discovery, they have gained a reputation as being pan assay interference compounds (PAINS) and frequent hitters in screening campaigns. Rhodanine-based compounds are also aggregators that can non-specifically interact with target proteins as well as Michael acceptors and interfere photometrically in biological assays due to their color.Areas covered: The authors review the recently reported biological activities of rhodanine-based compounds. Furthermore, the article provides details of their synthesis and occurrence in compound libraries through high-throughput screening (HTS) and virtual high-throughput screening (VHTS). Additionally, the authors provide the reader with possible mechanisms of non-specific target modulation, analysis of the crystal structures of enzyme–rhodanine complexes and a comparison of rhodanine and thiazolidine-2,4-dione moieties.Expert opinion: The biological activity of compounds possessing a rhodanine moiety should be considered very critically despite the convincing data obtained in biological assays. In addition to the lack of selectivity, unusual structure–activity relationship profiles and safety and specificity problems mean that rhodanines are generally not optimizable.

  • Research Article
  • Cite Count Icon 37
  • 10.1016/j.drudis.2021.02.003
Benchmarking the mechanisms of frequent hitters: limitation of PAINS alerts
  • Feb 10, 2021
  • Drug Discovery Today
  • Zi-Yi Yang + 6 more

Benchmarking the mechanisms of frequent hitters: limitation of PAINS alerts

  • Research Article
  • Cite Count Icon 62
  • 10.1016/j.chembiol.2021.01.021
Nuisance compounds in cellular assays
  • Feb 15, 2021
  • Cell chemical biology
  • Jayme L Dahlin + 15 more

Nuisance compounds in cellular assays

  • Research Article
  • Cite Count Icon 67
  • 10.1177/2472555218768497
Nuisance Compounds, PAINS Filters, and Dark Chemical Matter in the GSK HTS Collection
  • Jul 1, 2018
  • SLAS Discovery
  • Subhas J Chakravorty + 10 more

Nuisance Compounds, PAINS Filters, and Dark Chemical Matter in the GSK HTS Collection

  • PDF Download Icon
  • Research Article
  • Cite Count Icon 10
  • 10.1038/s41598-020-74139-0
Statistical models for identifying frequent hitters in high throughput screening
  • Oct 14, 2020
  • Scientific Reports
  • Samuel Goodwin + 2 more

High throughput screening (HTS) interrogates compound libraries to find those that are “active” in an assay. To better understand compound behavior in HTS, we assessed an existing binomial survivor function (BSF) model of “frequent hitters” using 872 publicly available HTS data sets. We found large numbers of “infrequent hitters” using this model leading us to reject the BSF for identifying “frequent hitters.” As alternatives, we investigated generalized logistic, gamma, and negative binomial distributions as models for compound behavior. The gamma model reduced the proportion of both frequent and infrequent hitters relative to the BSF. Within this data set, conclusions about individual compound behavior were limited by the number of times individual compounds were tested (1–1613 times) and disproportionate testing of some compounds. Specifically, most tests (78%) were on a 309,847-compound subset (17.6% of compounds) each tested ≥ 300 times. We concluded that the disproportionate retesting of some compounds represents compound repurposing at scale rather than drug discovery. The approach to drug discovery represented by these 872 data sets characterizes the assays well by challenging them with many compounds while each compound is characterized poorly with a single assay. Aggregating the testing information from each compound across the multiple screens yielded a continuum with no clear boundary between normal and frequent hitting compounds.

  • Research Article
  • Cite Count Icon 8
  • 10.2196/43758
Design, Development, and Evaluation of an Automated Solution for Electronic Information Exchange Between Acute and Long-term Postacute Care Facilities: Design Science Research.
  • Feb 17, 2023
  • JMIR Formative Research
  • Madhu Gottumukkala

Information exchange is essential for transitioning high-quality care between care settings. Inadequate or delayed information exchange can result in medication errors, missed test results, considerable delays in care, and even readmissions. Unfortunately, long-term and postacute care facilities often lag behind other health care facilities in adopting health information technologies, increasing difficulty in facilitating care transitions through electronic information exchange. The research gap is most evident when considering the implications of the inability to electronically transfer patients' health records between these facilities. This study aimed to design and evaluate an open standards-based interoperability solution that facilitates seamless bidirectional information exchange between acute care and long-term and postacute care facilities using 2 vendor electronic health record (EHR) systems. Using the design science research methodology, we designed an interoperability solution that improves the bidirectional information exchange between acute care and long-term care (LTC) facilities using different EHR systems. Different approaches were applied in the study with a focus on the relevance cycle, including eliciting detailed requirements from stakeholders in the health system who understand the complex data formats, constraints, and workflows associated with transferring patient records between 2 different EHR systems. We performed literature reviews and sought experts in the health care industry from different organizations with a focus on the rigor cycle to identify the components relevant to the interoperability solution. The design cycle focused on iterating between the core activities of implementing and evaluating the proposed artifact. The artifact was evaluated at a health care organization with a combined footprint of acute and postacute care operations using 2 different EHR systems. The resulting interoperability solution offered integrations with source systems and was proven to facilitate bidirectional information exchange for patients transferring between an acute care facility using an Epic EHR system and an LTC facility using a PointClickCare EHR system. This solution serves as a proof of concept for bidirectional data exchange between Epic and PointClickCare for medications, yet the solution is designed to expand to additional data elements such as allergies, problem lists, and diagnoses. Historically, the interoperability topic has centered on hospital-to-hospital data exchange, making it more challenging to evaluate the efficacy of data exchange between other care settings. In acute and LTC settings, there are differences in patients' needs and delivery of care workflows that are distinctly unique. In addition, the health care system's components that offer long-term and acute care in the United States have evolved independently and separately. This study demonstrates that the interoperability solution improves the information exchange between acute and LTC facilities by simplifying data transfer, eliminating manual processes, and reducing data discrepancies using a design science research methodology.

  • Research Article
  • Cite Count Icon 3
  • 10.21577/0103-5053.20240091
The Thin Line between Promiscuous and Privileged Structures in Medicinal Chemistry
  • Jan 1, 2024
  • Journal of the Brazilian Chemical Society
  • Thayná Bibiano + 3 more

The concept of privileged structures in medicinal chemistry refers to commonly found substructures in approved drugs and lead small molecules, presenting a broad profile of pharmacological action, i.e., participating in the recognition by several classes of pharmacological targets or by modulating physicochemical properties. Privileged structures are also related to nontoxic effects, which make their use in the design of new drug candidates very attractive. In contrast, another concept also refers to structures or substructures capable of presenting a pleiotropic profile for several pharmacological targets, referred to as promiscuous compounds. Worth mentioning, more recently, a great majority of promiscuous compounds have been classified as Pan Assay Interference Compounds (PAINS). In its great majority, PAINS are electrophilic in nature, capable of covalently reacting indiscriminately with several pharmacological targets, and are associated with specific substructures, which are, in turn, used as an exclusion filter during screening campaigns. This work aims to critically discuss the thin line that separates the two concepts, clarifying their differences from a molecular and pharmacological point of view. Moreover, special considerations regarding PAINS and exclusion filters will be made.

  • Book Chapter
  • Cite Count Icon 2
  • 10.1039/9781782622246-00214
Chapter 8. Rhodanine
  • Jan 1, 2015
  • Tihomir Tomašic + 1 more

Rhodanine has been associated with a wide variety of pharmacological activities and has been the target of extensive scientific debate in the past few years. Attention has been focused on the suggested problematic behaviour of compounds containing the rhodanine moiety, following the description by Baell et al. of rhodanines as frequent hitters or pan assay interference compounds (PAINS). The authors suggested, therefore, that rhodanines should be excluded from the compound libraries used for biomolecular screening and provided filters for eliminating such compounds using computational tools. Interest in rhodanines, nevertheless, still remains, with reports of their novel biological activities that highlight the rhodanine moiety as a privileged scaffold in the search for new biologically active compounds. This chapter describes the physicochemical properties and reactivities of rhodanines, followed by discussion about their biological activities, particularly their antibacterial, antiviral and anticancer activities. Finally, we highlight examples of rhodanine-based compounds progressing to clinical studies and use in therapy.

Save Icon
Up Arrow
Open/Close
Notes

Save Important notes in documents

Highlight text to save as a note, or write notes directly

You can also access these Documents in Paperpal, our AI writing tool

Powered by our AI Writing Assistant