Machine learning random forest for predicting oncosomatic variant NGS analysis

Eric Pellegrino,Nathalie Beaufils,Isabelle Nanni,Antoine Carlioz,Philippe Metellus,Coralie Jacques,L’Houcine Ouafik

doi:10.1038/s41598-021-01253-y

Abstract

Since 2017, we have used IonTorrent NGS platform in our hospital to diagnose and treat cancer. Analyzing variants at each run requires considerable time, and we are still struggling with some variants that appear correct on the metrics at first, but are found to be negative upon further investigation. Can any machine learning algorithm (ML) help us classify NGS variants? This has led us to investigate which ML can fit our NGS data and to develop a tool that can be routinely implemented to help biologists. Currently, one of the greatest challenges in medicine is processing a significant quantity of data. This is particularly true in molecular biology with the advantage of next-generation sequencing (NGS) for profiling and identifying molecular tumors and their treatment. In addition to bioinformatics pipelines, artificial intelligence (AI) can be valuable in helping to analyze mutation variants. Generating sequencing data from patient DNA samples has become easy to perform in clinical trials. However, analyzing the massive quantities of genomic or transcriptomic data and extracting the key biomarkers associated with a clinical response to a specific therapy requires a formidable combination of scientific expertise, biomolecular skills and a panel of bioinformatic and biostatistic tools, in which artificial intelligence is now successful in developing future routine diagnostics. However, cancer genome complexity and technical artifacts make identifying real variants challenging. We present a machine learning method for classifying pathogenic single nucleotide variants (SNVs), single nucleotide polymorphisms (SNPs), multiple nucleotide variants (MNVs), insertions, and deletions detected by NGS from different types of tumor specimens, such as: colorectal, melanoma, lung and glioma cancer. We compared our NGS data to different machine learning algorithms using the k-fold cross-validation method and to neural networks (deep learning) to measure the performance of the different ML algorithms and determine which one is a valid model for confirming NGS variant calls in cancer diagnosis. We trained our machine learning with 70% of our data samples, extracted from our local database (our data structure had 7 parameters: chromosome, position, exon, variant allele frequency, minor allele frequency, coverage and protein description) and validated it with the 30% remaining data. The model offering the best accuracy was chosen and implemented in the NGS analysis routine. Artificial intelligence was developed with the R script language version 3.6.0. We trained our model on 70% of 102,011 variants. Our best error rate (0.22%) was found with random forest machine learning (ntree = 500 and mtry = 4), with an AUC of 0.99. Neural networks achieved some good scores. The final trained model with the neural network achieved an accuracy of 98% and an ROC-AUC of 0.99 with validation data. We tested our RF model to interpret more than 2000 variants from our NGS database: 20 variants were misclassified (error rate < 1%). The errors were nomenclature problems and false positives. After adding false positives to our training database and implementing our RF model routinely, our error rate was always < 0.5%. The RF model shows excellent results for oncosomatic NGS interpretation and can easily be implemented in other molecular biology laboratories. AI is becoming increasingly important in molecular biomedical analysis and can be very helpful in processing medical data. Neural networks show a good capacity in variant classification, and in the future, they may be useful in predicting more complex variants.

Highlights

Tumor molecular profiles have recently become a major element in diagnosis, classification, and therapeutic management
We identify feedforward neural networks known as multilayer perceptrons (MLPs) where all arrows go in the same direction of the output and recurrent neural networks (RNNs) which might have a loop and tend to be much harder to train
Random forest (RF) is an machine learning (ML) algorithm that is defined by an ensemble of learning methods for classification, regression and other tasks that operate by constructing a multitude of decision trees at the training phase and outputting the classes or mean prediction of the individual trees

Summary

Introduction

Tumor molecular profiles have recently become a major element in diagnosis, classification, and therapeutic management. Data analysis of sequencing requires the use of bioinformatics pipelines to align the sequencing results on the human genome reference (hg19); and to filter mapped reads against the reference genome to perform variant analysis, including variant calling and predicting the effects produced by found variants on genes After these steps, biologists can begin the interpretation for each patient; they have to check a list of quality parameters to determine the mutation pathogenicity, by establishing the influence these mutations have on the protein. MLP is characterized by several layers of input nodes connected as a directed graph between the input and the output layers They use backpropagation to train the network and are widely used for solving problems that require supervised learning as well as research into computational neuroscience and parallel distributed processing. We validate the RF and ANN (artificial neural network) model by comparing the RF and MLP results to biologist decisions

Objectives

Methods

Results

Discussion

Conclusion

Talk to us

Join us for a 30 min session where you can share your feedback and ask us any queries you have

Schedule a call

Journal: Scientific Reports	Publication Date: Nov 8, 2021
Citations: 39	License type: open-access

R Discovery Prime

R Discovery Prime

Machine learning random forest for predicting oncosomatic variant NGS analysis

Abstract

Highlights

Summary

Talk to us

Similar Papers

More From: Scientific Reports

Lead the way for us

Similar Papers

Artificial intelligence: Friend or foe?
Anusch Yazdani ... Sam Costa
Australian and New Zealand Journal of Obstetrics and Gynaecology | VOL. 63
Anusch Yazdani, et. al.Anusch Yazdani ... Sam Costa
01 Apr 2023
Australian and New Zealand Journal of Obstetrics and Gynaecology | VOL. 63

Big data phenotyping in rare diseases: some ethical issues
Nina Hallowell ... Christoffer Nellåker
Genetics in Medicine | VOL. 21
Nina Hallowell, et. al.Nina Hallowell ... Christoffer Nellåker
01 Feb 2019
Genetics in Medicine | VOL. 21

Availability of Evidence for Predictive Machine Learning Algorithms in Primary Care
Margot M Rakers ... Hendrikus J A Van Os
JAMA Network Open | VOL. 7
Margot M Rakers, et. al.Margot M Rakers ... Hendrikus J A Van Os
03 Sep 2024
JAMA Network Open | VOL. 7

CORR Synthesis: When Should the Orthopaedic Surgeon Use Artificial Intelligence, Machine Learning, and Deep Learning?
Michael P Murphy ... Nicholas M Brown
Clinical orthopaedics and related research | VOL. 479
Michael P Murphy, et. al.Michael P Murphy ... Nicholas M Brown
17 Feb 2021
Clinical orthopaedics and related research | VOL. 479

Editage

Paperpal

R Discovery

Mind the Graph

R Discovery Prime

R Discovery Prime

Machine learning random forest for predicting oncosomatic variant NGS analysis

Abstract

Highlights

Summary

Talk to us

Similar Papers

More From: Scientific Reports