Properties of protein drug target classes.

Simon C Bull,Andrew J Doig

doi:10.1371/journal.pone.0117955

Simon C Bull, Andrew J Doig

Open Access

https://doi.org/10.1371/journal.pone.0117955

Copy DOI

Abstract

Accurate identification of drug targets is a crucial part of any drug development program. We mined the human proteome to discover properties of proteins that may be important in determining their suitability for pharmaceutical modulation. Data was gathered concerning each protein’s sequence, post-translational modifications, secondary structure, germline variants, expression profile and drug target status. The data was then analysed to determine features for which the target and non-target proteins had significantly different values. This analysis was repeated for subsets of the proteome consisting of all G-protein coupled receptors, ion channels, kinases and proteases, as well as proteins that are implicated in cancer. Machine learning was used to quantify the proteins in each dataset in terms of their potential to serve as a drug target. This was accomplished by first inducing a random forest that could distinguish between its targets and non-targets, and then using the random forest to quantify the drug target likeness of the non-targets. The properties that can best differentiate targets from non-targets were primarily those that are directly related to a protein’s sequence (e.g. secondary structure). Germline variants, expression levels and interactions between proteins had minimal discriminative power. Overall, the best indicators of drug target likeness were found to be the proteins’ hydrophobicities, in vivo half-lives, propensity for being membrane bound and the fraction of non-polar amino acids in their sequences. In terms of predicting potential targets, datasets of proteases, ion channels and cancer proteins were able to induce random forests that were highly capable of distinguishing between targets and non-targets. The non-target proteins predicted to be targets by these random forests comprise the set of the most suitable potential future drug targets, and should therefore be prioritised when building a drug development programme.

Highlights

The vast majority of the targets of approved drugs are proteins [1,2]
As a lower sequence identity threshold causes there to be a greater difference between the original dataset and the non-redundant one generated from it, using a range of thresholds enables classifier capability to be evaluated when the redundancy removal has different levels of influence on the dataset used for training
This enables a Random forests (RFs) induced using a non-redundant dataset to be evaluated in terms of its capability of generalising to the entire dataset, and allows the loss of information about the distribution of the proteins in the feature space, caused by the redundancy removal, to be assessed

Summary

Introduction

The vast majority of the targets of approved drugs are proteins [1,2]. Knowledge of which proteins are the targets of approved drugs enables the division of the human proteome into two classes: approved drug targets and non-targets. A protein is an approved drug target if it is the target of an approved drug, and a non-target otherwise. In order for a protein to have any potential as a drug target it must be druggable. A druggable protein is one that possesses folds that favour interactions with small drug-like molecules, PLOS ONE | DOI:10.1371/journal.pone.0117955. A druggable protein is one that possesses folds that favour interactions with small drug-like molecules, PLOS ONE | DOI:10.1371/journal.pone.0117955 March 30, 2015

Methods

Results

Discussion

Conclusion

Talk to us

Join us for a 30 min session where you can share your feedback and ask us any queries you have

Schedule a call

Journal: PloS one	Publication Date: Mar 30, 2015
Citations: 109	License type: CC BY 4.0

R Discovery Prime

R Discovery Prime

Properties of protein drug target classes.

Abstract

Highlights

Summary

Talk to us

Similar Papers

More From: PloS one

Lead the way for us

Similar Papers

In silico re-identification of properties of drug target proteins
Baeksoo Kim ... Chungoo Park
BMC Bioinformatics | VOL. 18
Baeksoo Kim, et. al.Baeksoo Kim ... Chungoo Park
01 May 2017
BMC Bioinformatics | VOL. 18

Predicting drug-target interaction based on sequence and structure information
Wei Lan ... Yi Pan
IFAC-PapersOnLine | VOL. 48
Wei Lan, et. al.Wei Lan ... Yi Pan
01 Jan 2015
IFAC-PapersOnLine | VOL. 48

ActiveDriverDB: Interpreting Genetic Variation in Human and Cancer Genomes Using Post-translational Modification Sites and Signaling Networks (2021 Update).
Michal Krassowski ... Miles W Mee
Frontiers in Cell and Developmental Biology | VOL. 9
Michal Krassowski, et. al.Michal Krassowski ... Miles W Mee
23 Mar 2021
Frontiers in Cell and Developmental Biology | VOL. 9

Efficient Data Mining Algorithms for Screening Potential Proteins of Drug Target
Qi Wang ... Jincai Huang
Mathematical Problems in Engineering | VOL. 2017
Qi Wang, et. al.Qi Wang ... Jincai Huang
01 Jan 2017
Mathematical Problems in Engineering | VOL. 2017

Editage

Paperpal

R Discovery

Mind the Graph

R Discovery Prime

R Discovery Prime

Properties of protein drug target classes.

Abstract

Highlights

Summary

Talk to us

Similar Papers

More From: PloS one