Identifying Cancer Drivers Using DRIVE: A Feature-Based Machine Learning Model for a Pan-Cancer Assessment of Somatic Missense Mutations.

Ionut Dragomir,Nirmesh Patel,Adnan Akbar,Gianmarco Contino,John W Cassidy,Harry W Clifford

doi:10.3390/cancers13112779

Abstract

Simple SummaryGenes dictate the grounds of life by comprising molecular bases which encode proteins. A mutation represents a gene modification that may influence the protein function. Cancer occurs when the mutation triggers uncontrolled cellular growth. Judging by the cancer expansion, mutations labelled as drivers confer a growth advantage, while passengers do not contribute to this augmentation. The aim of this study is methodological, which assesses the usefulness of a classification method for distinguishing between driver and passenger mutations. Based on 51 molecular characteristics of mutations and genes, including 3 novel features, multiple machine learning algorithms were used to determine whether these characteristics biologically represent the driver mutations and how they impact the classification procedure. To test the ability of the present methodology, the same steps were applied to an independent dataset. The results showed that both gene and mutation level characteristics are representative of the driver mutations, and the proposed approach achieved more than 80% accuracy in finding the true type of mutation. The evidence suggests that machine learning methods can be used to gain knowledge from mutational data seeking to deliver more targeted cancer treatment.Sporadic cancer develops from the accrual of somatic mutations. Out of all small-scale somatic aberrations in coding regions, 95% are base substitutions, with 90% being missense mutations. While multiple studies focused on the importance of this mutation type, a machine learning method based on the number of protein–protein interactions (PPIs) has not been fully explored. This study aims to develop an improved computational method for driver identification, validation and evaluation (DRIVE), which is compared to other methods for assessing its performance. DRIVE aims at distinguishing between driver and passenger mutations using a feature-based learning approach comprising two levels of biological classification for a pan-cancer assessment of somatic mutations. Gene-level features include the maximum number of protein–protein interactions, the biological process and the type of post-translational modifications (PTMs) while mutation-level features are based on pathogenicity scores. Multiple supervised classification algorithms were trained on Genomics Evidence Neoplasia Information Exchange (GENIE) project data and then tested on an independent dataset from The Cancer Genome Atlas (TCGA) study. Finally, the most powerful classifier using DRIVE was evaluated on a benchmark dataset, which showed a better overall performance compared to other state-of-the-art methodologies, however, considerable care must be taken due to the reduced size of the dataset. DRIVE outlines the outstanding potential that multiple levels of a feature-based learning model will play in the future of oncology-based precision medicine.

Talk to us

Join us for a 30 min session where you can share your feedback and ask us any queries you have

Schedule a call

Journal: Cancers	Publication Date: Jun 3, 2021
Citations: 4	License type: CC BY 4.0

R Discovery Prime

R Discovery Prime

Identifying Cancer Drivers Using DRIVE: A Feature-Based Machine Learning Model for a Pan-Cancer Assessment of Somatic Missense Mutations.

Abstract

Talk to us

Similar Papers

More From: Cancers

Lead the way for us

Similar Papers

Abstract 1176: Leveraging the GENIE dataset to distinguish somatic cancer drivers from passenger events in routine oncology practice
Philip A Beer ... Andrew V Biankin
Cancer Research | VOL. 82
Philip A Beer, et. al.Philip A Beer ... Andrew V Biankin
15 Jun 2022
Cancer Research | VOL. 82

Abstract 2183: Interrogating the molecular profile of colorectal cancer: detection of clinically actionable alterations in Hispanics
Ingrid M Montes-Rodriguez ... Noridza Rivera
Cancer Research | VOL. 82
Ingrid M Montes-Rodriguez, et. al.Ingrid M Montes-Rodriguez ... Noridza Rivera
15 Jun 2022
Cancer Research | VOL. 82

Abstract P3-06-20: Druggable genomic landscape of primary and metastatic breast cancers
Sq Sun
Cancer Research | VOL. 79
Sq SunSq Sun
15 Feb 2019
Abstract P3-06-20: Druggable genomic landscape of primary and metastatic breast cancers
Sq Sun

Getting Data Sharing Right to Help Fulfill the Promise of Cancer Genomics
Neil Savage
Cell | VOL. 168
Neil SavageNeil Savage
01 Feb 2017
Cell | VOL. 168

Editage

Paperpal

R Discovery

Mind the Graph

R Discovery Prime

R Discovery Prime

Identifying Cancer Drivers Using DRIVE: A Feature-Based Machine Learning Model for a Pan-Cancer Assessment of Somatic Missense Mutations.

Abstract

Talk to us

Similar Papers

More From: Cancers