A decision-theoretic approach to the evaluation of machine learning algorithms in computational drug discovery.

Oliver P Watson,Aimee R Taylor,Isidro Cortes-Ciriano,James A Watson

doi:10.1093/bioinformatics/btz293

Abstract

MotivationArtificial intelligence, trained via machine learning (e.g. neural nets, random forests) or computational statistical algorithms (e.g. support vector machines, ridge regression), holds much promise for the improvement of small-molecule drug discovery. However, small-molecule structure-activity data are high dimensional with low signal-to-noise ratios and proper validation of predictive methods is difficult. It is poorly understood which, if any, of the currently available machine learning algorithms will best predict new candidate drugs.ResultsThe quantile-activity bootstrap is proposed as a new model validation framework using quantile splits on the activity distribution function to construct training and testing sets. In addition, we propose two novel rank-based loss functions which penalize only the out-of-sample predicted ranks of high-activity molecules. The combination of these methods was used to assess the performance of neural nets, random forests, support vector machines (regression) and ridge regression applied to 25 diverse high-quality structure-activity datasets publicly available on ChEMBL. Model validation based on random partitioning of available data favours models that overfit and ‘memorize’ the training set, namely random forests and deep neural nets. Partitioning based on quantiles of the activity distribution correctly penalizes extrapolation of models onto structurally different molecules outside of the training data. Simpler, traditional statistical methods such as ridge regression can outperform state-of-the-art machine learning methods in this setting. In addition, our new rank-based loss functions give considerably different results from mean squared error highlighting the necessity to define model optimality with respect to the decision task at hand.Availability and implementationAll software and data are available as Jupyter notebooks found at https://github.com/owatson/QuantileBootstrap.Supplementary information Supplementary data are available at Bioinformatics online.

Highlights

Empirical methodologies guide a significant proportion of earlystage small-molecule drug discovery (Cumming et al, 2013; Keiser et al, 2007; Lipinski, 2004)
This work concerns the objective evaluation of the predictive ability of the latter, namely statistical and machine learning regression models trained on molecular structure-activity data
Use of regression modelling is often known as quantitative structure-activity relationship modelling (QSAR) (Sliwoski et al, 2014; Van De Waterbeemd and Gifford, 2003), and many different model classes have been used: support vector machines (Burbidge et al, 2001), ridge regression (Nandi et al, 2007), neural nets (Ajay et al, 1998; Lenselink et al, 2017; Nandi et al, 2007; Sadowski and Kubinyi, 1998) and random forests (Svetnik et al, 2003), to name but a few

Summary

Introduction

Empirical methodologies guide a significant proportion of earlystage small-molecule drug discovery (Cumming et al, 2013; Keiser et al, 2007; Lipinski, 2004). The goal of these models is to characterize the relationship between a highdimensional binary vector representation of small molecules (known as a molecular fingerprint) and the corresponding target specific in vitro activities In this context, use of regression modelling is often known as quantitative structure-activity relationship modelling (QSAR) (Sliwoski et al, 2014; Van De Waterbeemd and Gifford, 2003), and many different model classes have been used: support vector machines (Burbidge et al, 2001), ridge regression (Nandi et al, 2007), neural nets (Ajay et al, 1998; Lenselink et al, 2017; Nandi et al, 2007; Sadowski and Kubinyi, 1998) and random forests (Svetnik et al, 2003), to name but a few. The success of these models is in part due to high-throughput screening experiments which produce large structure-activity datasets (order of magnitude 102–106 datapoints)

Methods

Results

Conclusion

Full Text

Paper version not known

Open DOI Link

Talk to us

Join us for a 30 min session where you can share your feedback and ask us any queries you have

Schedule a call

Journal: Bioinformatics	Publication Date: May 9, 2019
Citations: 18	License type: CC BY 4.0

R Discovery Prime

R Discovery Prime

A decision-theoretic approach to the evaluation of machine learning algorithms in computational drug discovery.

Abstract

Highlights

Summary

Talk to us

Similar Papers

More From: Bioinformatics

Lead the way for us

Similar Papers

Evaluation of machine learning algorithms using OCT and OCT angiography for the diagnosis of multiple sclerosis
María Pilar Rojas Lozano ... Pablo García Mesa
Acta Ophthalmologica | VOL. 102
María Pilar Rojas Lozano, et. al.María Pilar Rojas Lozano ... Pablo García Mesa
01 Jan 2024
Acta Ophthalmologica | VOL. 102

Predicting and identifying factors associated with undernutrition among children under five years in Ghana using machine learning algorithms.
Eric Komla Anku ... Henry Ofori Duah
PLOS ONE | VOL. 19
Eric Komla Anku, et. al.Eric Komla Anku ... Henry Ofori Duah
13 Feb 2024
PLOS ONE | VOL. 19

Machine learning for the prediction of problems in steel tube bending process
Volkan Görüş ... Mehmet Çevik
Engineering Applications of Artificial Intelligence | VOL. 133
Volkan Görüş, et. al.Volkan Görüş ... Mehmet Çevik
16 May 2024
Engineering Applications of Artificial Intelligence | VOL. 133

Use of Multiprognostic Index Domain Scores, Clinical Data, and Machine Learning to Improve 12-Month Mortality Risk Prediction in Older Hospitalized Patients: Prospective Cohort Study.
Richard John Woodman ... Alberto Pilotto
Journal of medical Internet research | VOL. 23
Richard John Woodman, et. al.Richard John Woodman ... Alberto Pilotto
21 Jun 2021
Journal of medical Internet research | VOL. 23

Editage

Paperpal

R Discovery

Mind the Graph

R Discovery Prime

R Discovery Prime

A decision-theoretic approach to the evaluation of machine learning algorithms in computational drug discovery.

Abstract

Highlights

Summary

Talk to us

Similar Papers

More From: Bioinformatics