Cancer genome landscapes.
Over the past decade, comprehensive sequencing efforts have revealed the genomic landscapes of common forms of human cancer. For most cancer types, this landscape consists of a small number of "mountains" (genes altered in a high percentage of tumors) and a much larger number of "hills" (genes altered infrequently). To date, these studies have revealed ~140 genes that, when altered by intragenic mutations, can promote or "drive" tumorigenesis. A typical tumor contains two to eight of these "driver gene" mutations; the remaining mutations are passengers that confer no selective growth advantage. Driver genes can be classified into 12 signaling pathways that regulate three core cellular processes: cell fate, cell survival, and genome maintenance. A better understanding of these pathways is one of the most pressing needs in basic cancer research. Even now, however, our knowledge of cancer genomes is sufficient to guide the development of more effective approaches for reducing cancer morbidity and mortality.
- Research Article
7
- 10.3390/cancers12030701
- Mar 16, 2020
- Cancers
Matched-targeted and immune checkpoint therapies have improved survival in cancer patients, but tumor heterogeneity contributes to drug resistance. Our study categorized gene mutations from next generation sequencing (NGS) into three core processes. This annotation helps decipher complex biologic interactions to guide therapy. We collected NGS data on 145 patients who have failed standard therapy (2016 to 2018). One hundred and forty two patients had data for tissue (Caris MI/X) and plasma cell-free circulating tumor DNA (Guardant360) platforms. The mutated genes were categorized into cell fate (CF), cell survival (CS), and genome maintenance (GM). Comparative analysis was performed for concordance and discordance, unclassified mutations, trends in TP53 alterations, and PD-L1 expression. Two gene mutation maps were generated to compare each NGS platform. Mutated genes predominantly matched to CS with concordance between Guardant360 (64.4%) and Caris (51.5%). TP53 alterations comprised a significant proportion of the mutation pool in Caris and Guardant360, 14.7% and 13.1%, respectively. Twenty-six potentially actionable gene alterations were detected from matching ctDNA to Caris unclassified alterations. The CS core cellular process was the most prevalent in our study population. Clinical trials are warranted to investigate biomarkers for the three core cellular processes in advanced cancer patients to define the next best therapies.
- Research Article
- 10.1158/1538-7445.am2014-3920
- Sep 30, 2014
- Cancer Research
Large-scale cancer genome programs have generated rich data of genetic abnormalities observed in thousands of clinical patient tumors, which provides a major opportunity to develop a landscape of molecular aberrations across tumor types and propose new therapeutic targets. Although human tumor cell lines have been important tools for the cancer researcher for decades, there is a lack of bio-functional validation of genetic alterations and systematic molecular characterization across commonly used cell lines in this genomic age. To meet the need of applying cancer genome knowledge to facilitate basic and translational cancer research, an updated and accurate knowledge of the cell lines in use, particularly from a molecular viewpoint, is critical for in vitro studies. Here, we show systematic molecular characterization and clustering of 90 authenticated human tumor cell lines, which represent the most common human cancer types found in the clinic, such as lung, breast, colon, pancreatic, skin cancer and so on. By next generation sequencing and molecular profiling, those tumor cell lines were fully analyzed to capture the driver gene mutations, DNA copy number variations, gene expressions, protein expressions and relevant cell signaling pathway activations. In addition to driver mutations such as BRAF V600, KRAS G12, PI3K E545 and EGFR T790, the gene copy number amplifications of AKT, FGFR, MET, ERBB2 and deletion of PTEN are presented in this work. Moreover, correlating with the verified genetic alternations, the endogenous basal levels of protein expression and signaling pathway activations within the cells were analyzed by western blot and immunofluorescence staining. A set of 6 wild type control cell lines that were derived from either normal tissues or tumors were characterized in parallel. Furthermore, live cell growth kinetics was continually monitored by label-free cell based assay. Finally, the cell lines are clustered into 10 genetic alteration panels according to driver genes, critical protein kinases, or key components of signaling pathways. Cell lines are essential models for progressing our understanding of the cancer genome, as they allow for the functional and biological validation of the genetic alterations proposed by NGS data. The genetic alteration tumor cell panels and their molecular signature profiles provide powerful tools to accelerate the discoveries in basic cancer research, compound screening, biomarker selection, pathway analysis, and targeted therapeutic development. Citation Format: Lysa A. Volpe, Anupreet Bal, John Foulke, Michael Jackson, Luping Chen, Fang Tian. Understanding the molecular nature of cancer cell lines. [abstract]. In: Proceedings of the 105th Annual Meeting of the American Association for Cancer Research; 2014 Apr 5-9; San Diego, CA. Philadelphia (PA): AACR; Cancer Res 2014;74(19 Suppl):Abstract nr 3920. doi:10.1158/1538-7445.AM2014-3920
- Research Article
5
- 10.7717/peerj-cs.133
- Oct 9, 2017
- PeerJ Computer Science
Cataloging mutated driver genes that confer a selective growth advantage for tumor cells from sporadic passenger mutations is a critical problem in cancer genomic research. Previous studies have reported that some driver genes are not highly frequently mutated and cannot be tested as statistically significant, which complicates the identification of driver genes. To address this issue, some existing approaches incorporate prior knowledge from an interactome to detect driver genes which may be dysregulated by interaction network context. However, altered operations of many pathways in cancer progression have been frequently observed, and prior knowledge from pathways is not exploited in the driver gene identification task. In this paper, we introduce a driver gene prioritization method called driver gene identification through pathway and interactome information (DGPathinter), which is based on knowledge-based matrix factorization model with prior knowledge from both interactome and pathways incorporated. When DGPathinter is applied on somatic mutation datasets of three types of cancers and evaluated by known driver genes, the prioritizing performances of DGPathinter are better than the existing interactome driven methods. The top ranked genes detected by DGPathinter are also significantly enriched for known driver genes. Moreover, most of the top ranked scored pathways given by DGPathinter are also cancer progression-associated pathways. These results suggest that DGPathinter is a useful tool to identify potential driver genes.
- Research Article
15
- 10.1038/s41598-021-04015-y
- Jan 7, 2022
- Scientific Reports
An emergent area of cancer genomics is the identification of driver genes. Driver genes confer a selective growth advantage to the cell. While several driver genes have been discovered, many remain undiscovered, especially those mutated at a low frequency across samples. This study defines new features and builds a pan-cancer model, cTaG, to identify new driver genes. The features capture the functional impact of the mutations as well as their recurrence across samples, which helps build a model unbiased to genes with low frequency. The model classifies genes into the functional categories of driver genes, tumour suppressor genes (TSGs) and oncogenes (OGs), having distinct mutation type profiles. We overcome overfitting and show that certain mutation types, such as nonsense mutations, are more important for classification. Further, cTaG was employed to identify tissue-specific driver genes. Some known cancer driver genes predicted by cTaG as TSGs with high probability are ARID1A, TP53, and RB1. In addition to these known genes, potential driver genes predicted are CD36, ZNF750 and ARHGAP35 as TSGs and TAB3 as an oncogene. Overall, our approach surmounts the issue of low recall and bias towards genes with high mutation rates and predicts potential new driver genes for further experimental screening. cTaG is available at https://github.com/RamanLab/cTaG.
- Research Article
59
- 10.1371/journal.pcbi.1007381
- Sep 30, 2019
- PLoS Computational Biology
Cancer driver genes, i.e., oncogenes and tumor suppressor genes, are involved in the acquisition of important functions in tumors, providing a selective growth advantage, allowing uncontrolled proliferation and avoiding apoptosis. It is therefore important to identify these driver genes, both for the fundamental understanding of cancer and to help finding new therapeutic targets or biomarkers. Although the most frequently mutated driver genes have been identified, it is believed that many more remain to be discovered, particularly for driver genes specific to some cancer types. In this paper, we propose a new computational method called LOTUS to predict new driver genes. LOTUS is a machine-learning based approach which allows to integrate various types of data in a versatile manner, including information about gene mutations and protein-protein interactions. In addition, LOTUS can predict cancer driver genes in a pan-cancer setting as well as for specific cancer types, using a multitask learning strategy to share information across cancer types. We empirically show that LOTUS outperforms five other state-of-the-art driver gene prediction methods, both in terms of intrinsic consistency and prediction accuracy, and provide predictions of new cancer genes across many cancer types.
- Research Article
- 10.1158/1538-7445.fbcr15-ia17
- Feb 1, 2016
- Cancer Research
Cancer drug resistance is almost inevitable in the majority of patients with advanced metastatic tumors. Intratumor heterogeneity, facilitating rapid tumor evolution, is one of the main causes of tumor adaptation and resistance to cancer therapies. In this talk I will explore how cancer genome sequencing data can shed light on diversity within tumors and the processes shaping cancer genome evolution over space and time. Harnessing both extensive multi-region and single-sequencing data we have investigated intra-tumor heterogeneity across 10 major cancer types, focusing on non-small cell lung cancer, clear cell renal cell carcinoma, and esophageal adenocarcinoma. I will discuss our findings shedding light on the extent to which mutations in known driver genes, and “actionable mutations,” linked to targeted therapies, are clonal or subclonal. Tumor phylogenies and whether patterns of tumor evolution can be deciphered in these cancer types will also be explored. Temporal dissection of mutational signatures across cancer types will be used to reveal the dynamics of mutational processes during tumor evolution. I will explore how APOBEC mediated-mutagenesis can be linked to branched tumor evolution, and the acquisition of subclonal driver events across cancer types. Finally, the challenges and clinical implications of intratumor heterogeneity will be discussed. Citation Format: Nicholas McGranahan. Deciphering cancer genome evolution in heterogenous tumors. [abstract]. In: Proceedings of the Fourth AACR International Conference on Frontiers in Basic Cancer Research; 2015 Oct 23-26; Philadelphia, PA. Philadelphia (PA): AACR; Cancer Res 2016;76(3 Suppl):Abstract nr IA17.
- Research Article
- 10.1158/1538-7445.am2014-2372
- Sep 30, 2014
- Cancer Research
Driver somatic mutations confer a selective growth advantage to tumors while passenger mutations do not. Genes carrying driver mutations are detected in whole-exome sequencing (WES) studies if their non-synonymous mutations are significantly more frequent than their silent mutations, adjusting for the heterogeneous mutation rates across genomic locations, subjects and mutation types. Typical tumor WES studies (∼100-300 tumor samples) lack statistical power, and thus can identify only small numbers of candidate driver genes and possibly miss the well-established ones. The best algorithm to detect driver genes to date is MutSigCV. For each target gene, MutSigCV identifies its “bagel”, a set of genes predicted to have similar background mutation rates to the target gene and use the “bagels” to estimate the background mutation rate. Based on the identified “bagels”, we developed a novel statistical method implementing a series of algorithms to improve the statistical power while maintaining the type-I error rate. First, we developed a statistical framework for testing the selective growth advantage of each mutation type and appropriately integrating evidence of growth advantage across mutation types to achieve robust power. We then extended the algorithm to consider spatial clustering information. This substantially improves the statistical power when non-silent mutations are clustered into a few short regions or nucleotides. We applied this novel method to TCGA melanoma WES including 271 tumor samples. Without modeling spatial clustering information, over 30 candidate driver genes, including BRAF, NRAS, PTEN, TP53, PPP6C and CDKN2A were detected controlling FDR=0.01. Modeling spatial information of somatic mutations detected additional rarely mutated genes, including DISP1 (P=1.3×10-20, involved in cellular proliferation and differentiation that leads to normal development of embryonic structures), IDH1 (P=1.0×10-8, previously identified in relation to gliobalstoma and acute myeloid leukemia), MAP2K1 (P=5×10-18, a member of the mitogen-activated protein kinases that also include BRAF, MEK and other important players in melanoma), INTS8 (P=7.7×10-15, an integrator of small nuclear RNA processing), and STK19 (P=1.2×10-10, encoding a serine/threonine kinase, previously identified in melanoma). Finally, we developed methods to simultaneously analyze multiple cancers to improve the power for genes mutated in more than one cancer type, which will be illustrated by analyzing 12 major cancers in TCGA. Our method will provide an important tool for improving driver gene mutation detection in cancer. Citation Format: Xing Hua, Teresa Maria Landi, Jianxin Shi. Detecting driver genes based on tumor whole-exome sequencing studies. [abstract]. In: Proceedings of the 105th Annual Meeting of the American Association for Cancer Research; 2014 Apr 5-9; San Diego, CA. Philadelphia (PA): AACR; Cancer Res 2014;74(19 Suppl):Abstract nr 2372. doi:10.1158/1538-7445.AM2014-2372
- Research Article
63
- 10.1016/j.cels.2019.04.005
- May 1, 2019
- Cell Systems
Simultaneous Integration of Multi-omics Data Improves the Identification of Cancer Driver Modules.
- Research Article
109
- 10.1182/blood-2016-02-699843
- Jun 30, 2016
- Blood
Candidate driver genes involved in genome maintenance and DNA repair in Sézary syndrome
- Research Article
23
- 10.1002/cam4.704
- Mar 19, 2016
- Cancer Medicine
The driver genes play critical roles for tumorigenesis, and the number of identified driver genes reached plateau. But how they act during different cancer development stages is lack of knowledge. We investigated 138 driver genes’ mutation changes across clinical stages using 3,477 cases in nine cancer types from the Cancer Genome Atlas (TCGA) and constructed their temporal order relationships. We also examined the codon changes for the widely mutated TP53 and PIK3CA in tumor stages. Combinations of one to three driver genes specifically dominated in each cancer. Across the clinical stages, we categorized three patterns for the behaviors of driver genes’ mutation changes in the nine cancer types: recurrently mutated in all the stages and triggering other mutations; certain mutations lost meanwhile other mutations emerged; mutations dominated across entire stages, while other mutations gradually appeared or disappeared. We observed different codon changes dominated in different stages and revealed mutations recurrently occurring on the hotspot regions of the coding sequence may be the core factor for driver genes’ tumorigenesis. Our results highlighted the dynamic changes of oncogenesis roles in different clinical stages and suggested different diagnostic decision making according to the clinical stages of patients.
- Research Article
1
- 10.1259/bjr.20190625
- Feb 14, 2020
- The British Journal of Radiology
Although various single genetic factors have been shown to affect radiosensitivity, high-throughput DNA sequencing analyses have revealed complex genomic landscapes in many cancer types. The aim of this study is to elucidate the association between accumulated alterations in driver and passenger genes and radiation therapy response. We used 59 human solid cancer cell lines derived from 11 organ sites. Radiation-induced cell death was measured using a standard colony-forming assay delivered as a single dose ranging from 0 to 12 Gy. Comprehensive genomic data for the cell lines were acquired from the Catalogue Of Somatic Mutations In Cancer v. 80. Random forest classifiers were constructed to predict radioresistant phenotypes using genomic features. The Cancer Genome Atlas data sets were used to evaluate the clinical impact of the genomic feature following radiotherapy. The 59 cancer cell lines harbored either nucleotide variations or copy number variations in a median of 157 genes per cell. Radiosensitivity of the cancer cells was correlated with neither the number of driver gene mutations nor the number of passenger gene mutations. However, the proportion of driver gene alterations to total gene alterations in gene sets selected from the Kyoto Encyclopedia Genes and Genomes predicted radioresistant cells with sensitivity of 85% and specificity of 73%. High probability of radioresistance predicted by the model was associated with worse overall survival following definitive radiotherapy in patients of The Cancer Genome Atlas data sets. Cellular radiosensitivity was associated with the proportion of driver to total gene alterations in the selected oncogenic pathways, which may be a biomarker candidate for response to radiation therapy. These findings suggest that accumulated alterations in not only driver genes but also passenger genes affect radiosensitivity.
- Research Article
231
- 10.1016/j.bone.2016.10.017
- Oct 17, 2016
- Bone
Molecular genetics of osteosarcoma
- Research Article
10
- 10.1038/s10038-018-0481-4
- Jun 15, 2018
- Journal of Human Genetics
Extensive sequencing efforts of cancer genomes such as The Cancer Genome Atlas (TCGA) have been undertaken to uncover bona fide cancer driver genes which has enhanced our understanding of cancer and revealed therapeutic targets. However, the number of driver gene mutations is bounded, indicating that there must be a point when further sequencing efforts will be excessive. We found that there was a significant positive correlation between sample size and identified driver gene mutations across 33 cancers sequenced by the TCGA, which is expected if additional sequencing is still leading to the identification of more driver genes. However, the rate of new cancer driver genes being discovered with larger samples is declining rapidly. Our analysis provides a general guide for determining which cancer types would likely benefit from additional sequencing efforts, particularly those with relatively high rates of cancer driver gene discovery. Our results argue that past strategies of indiscriminately sequencing as many specimens as possible for all cancer types is becoming inefficient. In addition, without significant investments into applying our knowledge of cancer genomes, we risk sequencing more cancer genomes for the sake of sequencing rather than meaningful patient benefit.
- Research Article
38
- 10.1016/j.neucom.2018.03.026
- Mar 20, 2018
- Neurocomputing
A novel unsupervised learning model for detecting driver genes from pan-cancer data through matrix tri-factorization framework with pairwise similarities constraints
- Research Article
8
- 10.3389/fgene.2022.854190
- May 10, 2022
- Frontiers in Genetics
The progression of tumorigenesis starts with a few mutational and structural driver events in the cell. Various cohort-based computational tools exist to identify driver genes but require multiple samples to identify less frequently mutated driver genes. Many studies use different methods to identify driver mutations/genes from mutations that have no impact on tumor progression; however, a small fraction of patients show no mutational events in any known driver genes. Current unsupervised methods map somatic and expression data onto a network to identify personalized driver genes based on changes in expression. Our method is the first machine learning model to classify genes as tumor suppressor gene (TSG), oncogene (OG), or neutral, thus assigning the functional impact of the gene in the patient. In this study, we develop a multi-omic approach, PIVOT (Personalized Identification of driVer OGs and TSGs), to train on experimentally or computationally validated mutational and structural driver events. Given the lack of any gold standards for the identification of personalized driver genes, we label the data using four strategies and, based on classification metrics, show gene-based labeling strategies perform best. We build different models using SNV, RNA, and multi-omic features to be used based on the data available. Our models trained on multi-omic data improved predictions compared with mutation and expression data, achieving an accuracy for BRCA, LUAD, and COAD datasets. We show network and expression-based features contribute the most to PIVOT. Our predictions on BRCA, COAD, and LUAD cancer types reveal commonly altered genes such as TP53 and PIK3CA, which are predicted drivers for multiple cancer types. Along with known driver genes, our models also identify new driver genes such as PRKCA, SOX9, and PSMD4. Our multi-omic model labels both CNV and mutations with a more considerable contribution by CNV alterations. While predicting labels for genes mutated in multiple samples, we also label rare driver events occurring in as few as one sample. We also identify genes with dual roles within the same cancer type. Overall, PIVOT labels personalized driver genes as TSGs and OGs and also identified rare driver genes.