Nuclear receptors (NRs) are important biological targets of endocrine-disrupting chemicals (EDCs). Identifying chemicals that can act as EDCs and modulate the function of NRs is difficult because of the time and cost of in vitro and in vivo screening to determine the potential hazards of the 100000s of chemicals that humans are exposed to. Hence, there is a need for computational approaches to prioritize chemicals for biological testing. Machine learning (ML) techniques are alternative methods that can quickly screen millions of chemicals and identify those that may be an EDC. Computational models of chemical binding to multiple NRs have begun to emerge. Recently, a Nuclear Receptor Activity (NuRA) dataset, describing experimentally derived small-molecule activity against various NRs has been created. We have used the NuRA dataset to develop an ensemble of ML-based models to predict the agonism, antagonism, binding and effector binding of small molecules to nine different human NRs. We defined the applicability domain of the ML models as a measure of Tanimoto similarity to the molecules in the training set, which enhanced the performance of the developed classifiers. We further developed a user-friendly web server named 'NR-ToxPred' to predict the binding of chemicals to the nine NRs using the best-performing models for each receptor. This web server is freely accessible at http://nr-toxpred.cchem.berkeley.edu. Users can upload individual chemicals using Simplified Molecular-Input Line-Entry System, CAS numbers or sketch the molecule in the provided space to predict the compound's activity against the different NRs and predict the binding mode for each.
Read full abstract