UALCAN: A Portal for Facilitating Tumor Subgroup Gene Expression and Survival Analyses
UALCAN: A Portal for Facilitating Tumor Subgroup Gene Expression and Survival Analyses
- Research Article
- 10.1158/1538-7445.am2019-2481
- Jul 1, 2019
- Cancer Research
Introduction: The Cancer Genome Atlas (TCGA) consortium performed high-throughput sequencing of thousands of tumor and normal tissues across 33 cancer types. This has led to molecular characterization of different cancers and many novel discoveries. Such diverse data offers an excellent opportunity to further address the questions associated with tumor heterogeneity. Systematic exploration of epigenetic features and non-coding gene expression could lead clinicians/cancer researchers towards unearthing new diagnostic biomarkers and therapeutic targets. Previously our group developed a cancer transcriptome web portal, UALCAN [http://ualcan.path.uab.edu, Google: UALCAN)] to study expression and survival profile of protein coding genes among TCGA cancers based on various subgroups and molecular subtypes of cancer. Although many data portals facilitating easy access and in-depth analysis of TCGA data exists, there is need for a user-friendly resource facilitating comprehensive DNA methylation profile and gene expression analysis of lncRNA and miRNA within tumor subgroups based on factors such as stage, tumor grade, age, sex, race and integration with the gene expression. Methods: TCGA Level 3 RNA-seq, miRNA-seq and Illumina Infinium HumanMethylation450K data were downloaded using TCGA-Assembler 2 tool and Genomic Data Commons (GDC). The data were processed, organized based on tumor subtypes using custom R scripts. DNA methylation and gene expression analysis was performed using in-house PERL scripts, while statistical significance was estimated using “Statistics::TTest” module. The analyses results were stored as flat file database and hosted on a web portal developed using Apache2.0, PERL-CGI and HighChartJS JavaScript libraries. Results: The user friendly web portal 1) allows users could obtain list of top over-/under-expressed miRNAs and lncRNAs in each of 33 TCGA cancer 2) facilitates analysis of promoter methylation level and gene expression for given set of miRNA/ lncRNAs in various tumor sub-groups based on individual cancer stages, tumor grade, race, body weight or other clinicopathologic features, 3) enables researchers to download high resolution graphics such as heatmap, boxplots, dotplots as analyses outputs. The data can be downloaded in multiple output formats for use by researchers. Conclusion: The current resource aids cancer researchers in identifying DNA methylation mediated epigenetic modification of gene expression in tumor subgroup specific manner. Furthermore, our platform will facilitate analyses of long non-coding RNAs and microRNAs, thereby help in discovery of novel biomarkers and better understanding of the tumor biology. Citation Format: Darshan Shimoga Chandrashekar, Chad J. Creighton, Israel Ponce-Rodriguez, Sooryanarayana Varambally. Developing analysis platform for pan-cancer study of DNA methylation, mirna and lncrna expression based on tumor subtypes using TCGA data [abstract]. In: Proceedings of the American Association for Cancer Research Annual Meeting 2019; 2019 Mar 29-Apr 3; Atlanta, GA. Philadelphia (PA): AACR; Cancer Res 2019;79(13 Suppl):Abstract nr 2481.
- Research Article
33
- 10.1016/j.celrep.2021.109873
- Oct 1, 2021
- Cell Reports
Pan-cancer analysis of non-coding transcripts reveals the prognostic onco-lncRNA HOXA10-AS in gliomas.
- Research Article
- 10.1158/1538-7445.am2020-2317
- Aug 13, 2020
- Cancer Research
Background. The current best predictor of lifetime risk for many forms of cancer is family history. Family history is transmitted through germ line DNA. Analysis of germ-line DNA should be able to predict lifetime risk at least as well family history. Genetic risk scores, used to predict lifetime risk from DNA, usually consist of linear combinations of SNPs. These scores ignore structural variations and epistatic effects (non-linear combinations). The objective of this study is to test whether germline DNA structural information along with machine learning algorithms, which can use non-linear combinations, can be used to predict whether a person will develop different types of cancer. Methods. We tested this objective using data compiled by The Cancer Genome Atlas (TCGA) project and the UK Biobank. The TCGA data consisted of information on DNA chromosome scale length variation extracted from the peripheral blood of 8,821 different patients, each of whom had one of 32 different types of cancer. The UK Biobank data consisted of about 1500 women diagnosed with breast cancer (cases) and about 5000 women who have no record of any type of cancer(controls). We characterized structural variation by a quantitative measure of chromosomal-scale length variation. Chromosomal-scale length variation is computed from array data. A person's DNA (chromosomes 1-22) can be characterized by a series of 22 numbers, each representing the log base 2 ratio of the chromosome's length compared to the average chromosome length. Using these 22 numbers for each person, we set up a machine learning classification problem to differentiate those people diagnosed with a particular form of cancer from those who have not been diagnosed with that cancer. We used the h2o libraries in R to test how well different machine-learning algorithm can classify with these datasets. Results. In two independent datasets, the Cancer Genome Atlas (TCGA) project and the UK Biobank, we could classify whether or not a patient had a certain cancer based solely on chromosomal scale length variation. We found that all 32 different types of cancer in the TCGA dataset tested could be predicted better than chance using structural variation data. Specifically, in the TCGA dataset we measured the area under the receiver operator curve, known as the AUC, for ovarian cancer in women (0.89), glioblastoma multiforme (0.86), breast cancer in women (0.75), and colon adenocarcinoma (0.79), with a 95% confidence interval width of less than 0.01 in each case. This method could predict 10% of glioblastoma multiforme cases with less than 1 in 10,000 false positives in the TCGA dataset. In the UK Biobank dataset, we could predict breast cancer in women with an AUC=0.83. Conclusion. Genetic risk scores based on structural variation can effectively predict whether a person will eventually develop may different types of cancer. Citation Format: Chris Toh, Charmeine Ko, James P. Brody. Genetic risk scores from structural variation predict life-time individual cancer risk [abstract]. In: Proceedings of the Annual Meeting of the American Association for Cancer Research 2020; 2020 Apr 27-28 and Jun 22-24. Philadelphia (PA): AACR; Cancer Res 2020;80(16 Suppl):Abstract nr 2317.
- Research Article
2352
- 10.1016/j.ccr.2010.03.017
- Apr 15, 2010
- Cancer Cell
Identification of a CpG Island Methylator Phenotype that Defines a Distinct Subgroup of Glioma
- Research Article
- 10.1158/1538-7445.sabcs19-p2-18-03
- Feb 14, 2020
- Cancer Research
Background: For a decade, The Cancer Genome Atlas (TCGA) program collected clinicopathologic annotation data along with multi-platform molecular profiles of more than 11,000 human tumors across 33 different cancer types. TCGA clinical data were analyzed and a standardized dataset named the TCGA Pan-Cancer Clinical Data Resource was created. However, TCGA treatment data are not systematically analyzed. Here we focus on TCGA primary breast cancer (TCGA-BC) treatment data to assess their completeness, regimen patterns, and association with clinicopathologic features. Method: 814 TCGA-BC patients with treatment data diagnosed from 2001 to 2013 were selected for this study. The treatment data were prepared and classified to be adjuvant Chemotherapy (CT), adjuvant Radiation Therapy (RT), and adjuvant Hormone Therapy (HT). An in-house Clinical Breast Care Project (CBCP) treatment dataset (n=1051 from 2001 to 2017), which were relatively complete, were used as a reference for data completeness. Multinomial logistic regression was used to analyze the associations between different therapies and clinical features. Results: There is no consistent treatment data difference between TCGA-BC and CBCP. In TCGA-BC, 68.7%, 64.4%, and 64.5% of patients received CT, RT, and HT respectively, whereas in CBCP, the corresponding percentages were 59.0%, 71.2%, and 77.1%. The percentages of patients receiving combined therapies are even more comparable between the two cohorts (data not shown due to many combinations). The associations between treatment and clinicopathologic features were analyzed using multinomial logistic regressions with age, race, menopausal status, AJCC stage, and PAM50 subtype as covariates. Several of these covariates significantly associate with the use of therapies. Compared to patients with age < 50, patients with age ≥ 65 were significantly more likely to receive HT, HT+RT, or RT than CT only (OR=11.34, 9.25, 4.82; p=0.001, 0.002, and 0.024 respectively); Compared to pre-menopausal patients, post-menopausal patients were significantly more likely to receive HT+RT than CT only (OR=4.12, p=0.034); Compared to patients with early stages (I, II), patients with advanced stages (III, IV) were more likely to receive CT+HT+RT (OR=1.92, p=0.038) or CT+RT (OR=2.38, p=0.010) than CT only, and less likely to receive CT+HT (OR=0.39, p=0.033) or HT (OR=0.30, p=0.008) than CT only. PAM50 subtype also significantly associates with the use of therapies. Compared to Luminal A patients, patients with Basal subtype were significantly less likely to receive CT+HT (OR=0.14, p=1.3e-05), CT+HT+RT (OR=0.06, p=3.0e-10), HT+RT (OR=0.01, p=9.2e-06), or HT (OR=0.02, p=3.8e-07) than CT alone. Patients with HER2+ subtype showed similar patterns (note targeted therapies were classified as CT). In addition, patients with Basal subtype were significantly more likely to receive CT+RT (OR=2.47, p=0.01) than CT only. Conclusion: TCGA-BC patient treatment data are relatively as complete as our in-house CBCP patient treatment data, enabling us to perform a preliminary analysis of the former for association with clinicopathologic features of the patients. The significant associations of age, menopausal status, AJCC stage, and PAM50 subtype with treatment regimens are consistent with clinical knowledge, suggesting potential validity of TCGA BC treatment data for research use. Disclaimer: The contents of this publication are the sole responsibility of the author(s) and do not necessarily reflect the views, opinions or policies of Uniformed Services University of the Health Sciences (USUHS), The Henry M. Jackson Foundation for the Advancement of Military Medicine, Inc., the Department of Defense (DoD), the Departments of the Army, Navy, or Air Force. Mention of trade names, commercial products, or organizations does not imply endorsement by the U.S. Government. Citation Format: Jianfang Liu, J. Leigh Fantacone-Campbell, Albert J. Kovatich, Brenda Deyarmin, Bradley J. Mostoller, Jeffrey A. Hooke, Hallgeir Rui, Craig D. Shriver, Hai Hu. Breast cancer treatment and association with clinicopathologic features in TCGA [abstract]. In: Proceedings of the 2019 San Antonio Breast Cancer Symposium; 2019 Dec 10-14; San Antonio, TX. Philadelphia (PA): AACR; Cancer Res 2020;80(4 Suppl):Abstract nr P2-18-03.
- Research Article
2
- 10.1158/1538-7445.am2024-6548
- Mar 22, 2024
- Cancer Research
The NCI's The Cancer Genome Atlas (TCGA) project profiled over 10,000 tumor samples over the course of 10 years. As different tissue-specific working groups reviewed all of the available data, these patient samples were separated into distinct molecular subtypes, and these clusters were reported in various marker papers. While these assignments provided invaluable information about the common patterns of molecular characteristics in different types of cancer there was no consistent methodology for assigning new samples to these defined molecular subtypes.The NCI's Tumor Molecular Pathology group was formulated to create machine learning-based models that could be applied to non-TCGA samples and determine their TCGA mapped subtypes. Five modeling systems, JADBio, SKGrid by the Oregon Health and Science University, CloudForest by the Institute of Systems Biology, AKLIMATE by University of California Santa Cruz and subSCOPE by BC Cancer’s Genome Sciences Centre, were trained to recognize TCGA subtypes using multi-omic measurements from gene expression, DNA methylation, miRNA expression, copy number, and somatic mutation calls. While the TCGA samples were profiled using multi-omic technologies, single platform and/or compact feature set models also were assessed for their ability to assign these classifications. Each machine learning system created predictive models for 106 subtypes from 26 cancer types using as few features as possible, with a maximum of 100 features allowed for scored models. A set of 411,706 models was developed, composed of results of each of the learning methods across the various omic platforms. Top models, both multi-omic and single platform, were selected for each cancer type. On average, models were able to achieve an overall weighted F1 score of 0.895 with 42 features. While the top models for each cancer type had an overall weighted F1 mean performance of 0.936 with a mean of 29 features, in 20 of the 26 cancer types models using only gene expression provided the best performance. Analysis of features selected by the models showed some known onco-drivers were selected by many models, but many times different models would utilize features of different genes with similar levels of performance. Network-level analysis revealed that many genes of these selected features operated within the same pathways.Transferability of these models to external datasets was tested, taking TCGA breast cancer trained models and applying them to AURORA and METABRIC datasets. Interestingly, despite the data platform difference between TCGA (RNAseq) and METABRIC (microarray), model performance saw only minimal degradation of F1 values in transfer. This set of models and the training dataset will provide new opportunities for researchers and translational scientists to connect new tumors to the subtypes seen in the TCGA cohorts. Citation Format: Kyle Ellrott, Chris K. Wong, Christina Yau, Mauro A. Castro, Jordan Lee, Brian Karlberg, Jasleen K. Grewal, Vincenzo Lagani, Bahar Tercan, Verena Friedl, Toshinori Hinoue, Vladislav Uzunangelov, Lindsay Westlake, Xavier Loinaz, Ina Felau, Peggy Wang, Anab Kemal, Samantha J. Caesar-Johnson, Ilya Shmulevich, Alexander J. Lazar, Ioannis Tsamardinos, Katherine A. Hoadley, The Cancer Genome Atlas Analysis Network, Gordon A. Robertson, Theo A. Knijnenburg, Christopher C. Benz, Joshua M. Stuart, Jean C. Zenklusen, Andrew D. Cherniack, Peter W. Laird. Leveraging compact feature sets for TCGA-based molecular subtype classification on new samples [abstract]. In: Proceedings of the American Association for Cancer Research Annual Meeting 2024; Part 1 (Regular Abstracts); 2024 Apr 5-10; San Diego, CA. Philadelphia (PA): AACR; Cancer Res 2024;84(6_Suppl):Abstract nr 6548.
- Research Article
34
- 10.1002/cjp2.271
- Apr 5, 2022
- The Journal of Pathology: Clinical Research
Homologous recombination deficiency (HRD) leads to DNA double‐strand breaks and can be exploited by the use of poly (ADP‐ribose) polymerase (PARP) inhibitors to induce synthetic lethality. Extending the original therapeutic concept, the role of HRD is currently being investigated in clinical trials testing immune checkpoint blockers alone or in combination with PARP inhibitors, but the relationship between HRD and immune cell context in cancer is incompletely understood. We analyzed the association between immune cell composition, gene expression, and HRD in 9,041 tumors of 32 solid cancer types from The Cancer Genome Atlas (TCGA). The numbers of genomic scars were quantified by the HRD sum score (HRDsum) including loss of heterozygosity, large‐scale state transitions, and telomeric allelic imbalance. The T‐cell inflamed gene expression profile correlated weakly, but significantly positively, with HRDsum across cancer types (ρ = 0.17). Within individual cancer types, a significantly positive correlation was observed only in breast cancer, ovarian cancer, and four other cancer types, but not in the remaining 26 cancer types. HRDsum and tumor mutational burden (TMB) correlated significantly positively across cancer types (ρ = 0.42) and within 18 cancer types. HRDsum and a proliferation metagene correlated significantly positively across cancer types (ρ = 0.52) and within 20 cancer types. Mismatch repair deficiency and HRD as well as proofreading deficiency showed a high level of exclusivity. High HRD scores were associated with an immunologically activated tumor microenvironment only in a minority of cancer types. Our data favor the combination of genetic markers, complex genomic markers (including HRDsum and TMB), and other molecular markers (including proliferation scores) for a precise and comprehensive read‐out of the tumor biology and an individually tailored treatment.
- Research Article
- 10.1158/1538-7445.transcagen-a1-32
- Nov 15, 2015
- Cancer Research
The Stanford-TCGA Portal (http://genomeportal.stanford.edu/pan-tcga) enables users to easily navigate cancer genomic/proteomic data of cancer patients with their clinical information from the Cancer Genome Atlas (TCGA) project. As of 2014 January, the TCGA processed thousands of samples from over 20 cancer types and the data is publically available. This huge data set can offer many valuable insights about cancer research. However, exploring data from TCGA remains a challenge, particularly for researchers and others who lack formal bioinformatics training. With the Stanford-TCGA Portal, anyone can use a personal device such as laptops or smart phones to explore the TCGA data in regards to clinically relevant questions. For example, “What genes are associated with advanced breast cancer?”, “What is the difference in the frequency of TP53 mutations between diffuse and intestinal gastric cancers?”, or “What is the frequency of copy number deletion of APC for samples with/without PIK3CA mutations?”. While several websites such as cBio portal or UCSC cancer genomics browser was developed in order to make TCGA data more accessible to a community, there are few interactive features that enable querying of clinically-relevant phenotypic associations with drivers and modifier genes. The Stanford-TCGA portal allows users to navigate TCGA data in three different ways; i) search for clinically relevant genes/micro RNAs (miRs)/proteins by names, cancer types or clinical parameters, ii) profile genomic/proteomic changes by clinical parameters in a cancer type, or iii) test two-hit hypothesis. Thus, one can address specific questions regarding the association of cancers drivers with clinical parameters and outcome. We identify clinically relevant genes/miRs/proteins by our rigorous statistical analysis, which has been previously published, from 17 cancer types with 19 clinical parameters such as clinical stage or smoking history. Basically, elastic-net method estimates an optimal multiple linear regularized regression of the clinical parameters on the space of genomic/proteomic features. This analysis identified the set of top gene predictors of each clinical parameter in each cancer type respectively. Users can easily access the lists of these identified genes through the Stanford-TCGA Portal. In summary, the Stanford-TCGA Portal enables the cancer research community and others to fully utilize TCGA data. It provides simple yet clinically relevant information for web and mobile interface, thus one can examine queries and test hypothesis regarding genomic/proteomic alterations in cancers from any time and place. This will be an important step towards translation of genomic/proteomic data into clinics. Citation Format: HoJoon Lee, Jennifer Palm, Hanlee Ji. The Stanford-TCGA portal: An interactive web/mobile interface for exploring the clinical phenotypic relevance of specific cancer drivers. [abstract]. In: Proceedings of the AACR Special Conference on Translation of the Cancer Genome; Feb 7-9, 2015; San Francisco, CA. Philadelphia (PA): AACR; Cancer Res 2015;75(22 Suppl 1):Abstract nr A1-32.
- Research Article
- 10.1158/1538-7445.am2013-3171
- Apr 15, 2013
- Cancer Research
Using data from The Cancer Genome Atlas (TCGA) project and HapMap consortium, we examined the hypothesis that existing copy number polymorphisms might predispose individuals to cancer and if different common CNVs bias toward different cancer types. This analysis was performed using a subset of the TCGA data, obtained from blood derived “normal” samples of cancer patients. All the data was generated using Affymetrix SNP 6.0 arrays (Affymetrix, Inc.) and processed using Nexus Copy Number ver. 6.1 (BioDiscovery, Inc.). The project contained 451 Glioblastoma (GBM), 353 Colon cancer, and 245 Skin cancer samples in addition to 211 HapMap samples which we used as reference since these individuals were supposed to be healthy. Using Fisher's Exact test, we compared each different cancer types to the HapMap samples to identify regions of copy number change with FDR corrected p-value of less than 0.05 and difference in frequency between each population of at least 10%. Many of the regions found in this process encompass regions identified as common CNV polymorphisms as reported in the Database of Genomic Variance (DGV). For example, a polymorphic region on 8p23.1 containing a number of genes from the Beta defensin family involved in defense response to bacterium were seen significantly gained in germline samples obtained from cancer patients as compared to the HapMap samples (11% in HapMap, 42% in Colon Cancer, 28% in GBM, and 32% in Skin cancer). Additionally, some interesting regions were identified that were not seen as polymorphic in the general population, did not show up in the HapMap samples or reported in DGV, but had high prevalence in the blood derived normal samples from cancer patients. Using gene enrichment analysis on genes discovered in these regions we identified a number of known cancer specific pathways that were highly enriched (e.g. KEGG Glioma and Melanoma pathways being statistically significantly enriched). Looking at the effected genes in these pathways we identified small focal and recurrent events that often only spanned a small portion of a gene covering a few exons. This included genes such as PIK3CA where a focal gain covering exons 10-14 were observed in none of the HapMap samples but in 21% of the germline of patients with Colon cancer, 9% of GBM and 14% of Skin cancer patients. A similar focal and recurrent gain was observed in KIT covering exons 2-7 with no gains in the HapMap samples and 55% in Colon, 15% in GBM, and 50% in Skin cancer patients. Similar observations were also made for other well-known cancer genes such as RAF1 and EGFR but to a lesser extent. We are in a process of comparing matched tumor/normal cases to further investigate possible evolution of these changes from the germline to the surrounding tissue (also available from the TCGA database) to the primary tumor samples. Citation Format: Soheil Shams, Raja Keshavan, Joshua D. Schiffman. Analysis of copy number and LOH in germline samples from different tumor types in The Cancer Genome Atlas (TCGA) project. [abstract]. In: Proceedings of the 104th Annual Meeting of the American Association for Cancer Research; 2013 Apr 6-10; Washington, DC. Philadelphia (PA): AACR; Cancer Res 2013;73(8 Suppl):Abstract nr 3171. doi:10.1158/1538-7445.AM2013-3171
- Research Article
2
- 10.1097/cm9.0000000000002657
- Apr 19, 2023
- Chinese Medical Journal
To the Editor: Acute myeloid leukemia (AML) is a heterogeneous disease characterized by proliferation of abnormal myeloblasts in the bone marrow. Traditional treatment, such as chemotherapy with or without hematopoietic stem cell transplantation (HSCT), has been used as the first-line therapy in AML. However, the clinical outcome of patients with AML varies greatly due to several factors, such as the tumor immune microenvironment (TME).[1] Immune checkpoint inhibitors (ICIs) have recently gained popularity as therapeutic options for relapsed AML or MRD-positive AML patients.[2] Based on previous studies, immune dysregulation plays an important role in AML relapse, and immune-related genes may provide prognostic information for AML patients. Thus, exploring the immune-related genetic prognostic index (IRGPI) to predict prognosis and therefore provide personalized guidance for ICI therapy is urgently needed. In this study, we developed an immune-related genetic prognostic index for AML that was able to predict overall survival (OS) and ICI immunotherapy benefit. We focused on all immune-related genes in AML transcriptome data and screened immune-related genes associated with prognosis by weighted gene coexpression network analysis (WGCNA) to construct an immune-related gene prognostic index (IRGPI). We then characterized the immune and molecular features of the IRGPI, examined its prognostic ability for AML patients, and compared it with tumor immune dysfunction and exclusion (TIDE) as well as the tumor inflammation signature (TIS).[3,4] The results showed that the IRGPI is a promising prognostic biomarker for AML patients to predict the survival and benefits of ICI immunotherapy. First, transcriptome and clinical information of the cohort The Cancer Genome Atlas (TCGA)-AML and all data for normal tissue samples from 336 whole blood samples in Genotype-Tissue Expression (GTEx) were retrieved via the UCSC Xena platform. RNA-seq data for 104 AML samples (GSE71014) and survival information were downloaded from the Gene Expression Omnibus (GEO). The clinical characteristics of 132 AML patients from TCGA-LAML used for survival analysis are listed in Supplementary Table 1, https://links.lww.com/CM9/B503. Lists of immune-related genes were downloaded from the ImmPort and InnateDB database. A total of 14,672 differentially expressed genes were obtained in differential expression analysis. By intersecting these genes with the lists of immune-related genes obtained from ImmPort and InnateDB, 908 differentially expressed immune-related genes were identified, of which 189 genes were upregulated and 719 downregulated in tumor samples compared with normal samples. WGCNA was carried out on candidate genes to obtain immune-related hub genes (n = 908), revealing the top 117 immune-related hub genes with a threshold degree of >20. Expression of 24 of these immune-related hub genes correlated closely with AML patient OS, as determined by Kaplan–Meier (K–M) analysis (P <0.001). To determine independent prognostic genes, multivariate Cox regression analysis for OS was performed among the 24 immune-related hub genes, and only seven genes (CLEC11A, IL1R2, IL1RL2, TRIM55, TREML2, CAMK2A, and BMP2) were found to significantly affect the OS of AML patients. Then, we constructed a prognostic index for all cancer samples calculated by the following formula: IRGPI = expression level of CLEC11A × (-0.59) + expression level of IL1R2 × 0.37 + expression level of IL1RL2 × 3.63 + expression level of TRIM55 × (4.98) + expression level of TREML2 × (1.25) + expression level of CAMK2A × (2.17) + expression level of BMP2 × (0.83). Taking the median IRGPI as the cut-off value, IRGPI-high patients had a lower OS than IRGPI-low patients (P <0.001) among AML samples from TCGA. Then, the role of the IRGPI was validated using the GSE71014 cohort. The patients in the IRGPI-low subgroup had a significantly better prognosis than those in the IRGPI-high subgroup (P = 0.044, data not shown), which was consistent with the result using the TCGA dataset and indicated the prognostic value of the IRGPI. Multivariate Cox regression analysis confirmed the IRGPI to be an independent prognostic factor after adjusting for other clinicopathologic factors (P <0.001). Next, we explored the relationship between the IRGPI score and CD274 expression as well as CTLA4 and LAG3. The IRGPI score correlated significantly positively with CD274 expression (r = 0.42, P <0.001) and slightly with CTLA4 and LAG3 expression (r =0.3, P <0.001; r = 0.27, P <0.001) (data not shown). To analyze the composition of immune cells in different IRGPI subgroups, we used the Cell-type Identification by Estimating Relative Subsets of RNA Transcripts (CIBERSORT) algorithm to evaluate differences in immune cells between the IRGPI-high and IRGPI-low groups. We found that monocytes and eosinophils were more abundant in the IRGPI-high subgroup but that resting mast cells were more abundant in the IRGPI-low subgroup. Further analysis showed that higher monocytes correlated with lower OS ability and that higher resting mast cells were associated with longer OS (data not shown). We then used TIDE to evaluate the potential clinical efficacy of immunotherapy in the two IRGPI groups. Previous studies have shown that a higher TIDE prediction score is related to a higher potential for immune evasion, indicating that such patients are less likely to benefit from ICI therapy.[3] Our results showed that the IRGPI-high group had a higher TIDE score than the IRGPI-low subgroup [Figure 1A], suggesting that IRGPI-low patients may benefit more from ICI therapy than IRGPI-high patients. Moreover, previous studies have shown that a higher TIDE prediction score is associated with a worse outcome. Therefore, IRGPI-low patients with a low TIDE score might have a better prognosis than IRGPI-high patients with a high TIDE score. In addition, the IRGPI-high group had a higher risk of T-cell exclusion and T-cell dysfunction than the IRGPI-low group [Figure 1B, C]. We also performed receiver operating characteristic (ROC) analysis of the IRGPI for OS at 1-year, 2-year, and 3-year follow-up, and Figure 1D shows a robust prediction ability at 3 years (AUC: 0.819). The area under the ROC curve for the IRGPI was higher than that for TIDE or TIS [Figure 1E]. Further analysis revealed a correlation between the IRGPI-high group and lower survival probability in the IMvigor210 cohort[5] [Figure 1F], suggesting a strong prognosis ability compared with TIDE and TIS in AML and a possible indicator for ICI therapy benefit.Figure 1: The benefit of ICI therapy in the two IRGPI subgroups. (A–C) TIDE, T-cell dysfunction, and exclusion score in different IRGPI subgroups. The score between the two IRGPI subgroups was compared through the Wilcoxon test (* P < 0.001, † P <0.01). (D) ROC analysis of IRGPI on overall survival at 12-, 24-, and 36-month follow-up in TCGA cohort. (E) ROC analysis of IRGPI, TIS, and TIDE on overall survival at 36-month follow-up in TCGA cohort. (F) Kaplan–Meier survival analysis of the IRGPI subgroups in the IMvigor210 cohort. ICI: Immune checkpoint inhibitor; IRGPI: Immune-related genetic prognostic index; ROC: Receiver operating characteristic; TCGA: The Cancer Genome Atlas; TIDE: Tumor immune dysfunction and exclusion.By using RNA-seq data of AML, we identified a prognostic biomarker, the IRGPI, for AML. The IRGPI proved to be a valid prognostic immune-related biomarker for AML, with better survival in IRGPI-low patients and worse survival in IRGPI-high patients in both TCGA and GEO cohorts.In our study, the composition of some immune cells differed between the two IRGPI subgroups: the IRGPI-high subgroup had higher levels of monocytes and neutrophils, whereas resting mast cells were significantly enriched in the IRGPI-low subgroup. It has been reported that monocytes can be driven by AML blasts to differentiate into the M2-like phenotype, which promotes a more immunosuppressive microenvironment and is averse to AML clearance.[6] This is in accordance with the result that higher monocytes correlated with lower OS in our study. Therefore, the phenotype and functions of neutrophils in AML should be carefully evaluated in the future. Previous studies have proven that biomarkers, such as TIDE and TIS, can predict patient response to immunotherapy. TIDE can be used to identify two immune escape mechanisms that induce T-cell dysfunction in tumors with high CTL invasion and prevent T-cell invasion in tumors with low CTL levels.[3] In addition, TIS provides both quantitative and qualitative information about the TME, and 18 genes that reflect an ongoing CD8 T-cell response have shown promising results in predicting response to anti-PD-1/PD-L1 agents.[4] However, both TIDE and TIS focus on the immune function of T cells. Since the composition of immune cell subtypes and expression of immunosuppressive molecules were different between the two IRGPI subgroups, the IRGPI might reflect different immune benefits from ICR therapy identified with TIDE. Furthermore, both markers relate to patient response to immunotherapy rather than patient survival time. Encouragingly, the predictive value of the IRGPI in our study was comparable to that of TIDE and TIS, and the IRGPI might be a better predictor of OS at longer follow-up times. We further confirmed the prognostic value of the IRGPI in predicting the survival of patients who receive ICI therapy. Moreover, as the IRGPI is composed of only seven genes, it is easier to implement than TIDE and TIS. In conclusion, the IRGPI is a promising immune-related prognostic biomarker that may help in distinguishing immune and molecular characteristics and predicting AML patient clinical outcomes. The IRGPI might serve as a prognostic indicator for response to ICI immunotherapy; however, our prognostic model was limited to data from public databases and lacked more prospective real-world data. Therefore, more studies are needed to verify our results. Second, as we only considered an index of immune-related genes, many prominent prognostic genes in AML might have been excluded. It is possible that the IRGPI can be applied jointly with other types of biomarkers to achieve higher prediction performance. Finally, it should be emphasized that the association between the risk score and immune activity has not yet been experimentally addressed. Funding This study was supported by grants from the National Natural Science Foundation of China (Nos. 81670166, 81870140, 82070184, and 81621001), Peking University People's Hospital Research and Development Funds (No. RDL2021-01), Beijing Life Oasis Public Service Center (No. CARTFR-01), and CSH Young Scholars and 3SBio Pharmaceutical joint research project (No. KYC2201001). Conflict of Interest None.
- Research Article
32
- 10.1126/scitranslmed.ads6335
- Sep 3, 2025
- Science translational medicine
In recent years, a growing number of publications have reported the presence of microbial species in human tumors and mixtures of microbes that appear to be highly specific to different cancer types. Our recent reanalysis of data from three cancer types revealed that technical errors have caused erroneous reports of numerous microbial species found in sequencing data from The Cancer Genome Atlas (TCGA) project. Here, we have expanded our analysis to cover all 5734 whole-genome sequencing (WGS) datasets currently available from TCGA, covering 25 distinct types of cancer. We analyzed the microbial content using updated computational methods and databases and compared our results to those from two major recent studies that focused on bacteria, viruses, and fungi in cancer. Our results expand upon and reinforce our recent findings, which show that the presence of microbes is far smaller than had been previously reported and that many species identified in TCGA data might not be present at all. As part of this expanded analysis and to help others avoid being misled by flawed data, we have released a dataset that contains detailed read counts for bacteria, viruses, archaea, and fungi detected in all 5734 TCGA samples, which can serve as a public reference for future investigations.
- Research Article
11
- 10.1101/2024.05.24.595788
- Aug 19, 2024
- bioRxiv
In recent years, a growing number of publications have reported the presence of microbial species in human tumors and of mixtures of microbes that appear to be highly specific to different cancer types. Our recent re-analysis of data from three cancer types revealed that technical errors have caused erroneous reports of numerous microbial species found in sequencing data from The Cancer Genome Atlas (TCGA) project. Here we have expanded our analysis to cover all 5,734 whole-genome sequencing (WGS) data sets currently available from TCGA, covering 25 distinct types of cancer. We analyzed the microbial content using updated computational methods and databases, and compared our results to those from two major recent studies that focused on bacteria, viruses, and fungi in cancer. Our results expand upon and reinforce our recent findings, which showed that the presence of microbes is far smaller than had been previously reported, and that many species identified in TCGA data are either not present at all, or are known contaminants rather than microbes residing within tumors. As part of this expanded analysis, and to help others avoid being misled by flawed data, we have released a dataset that contains detailed read counts for bacteria, viruses, archaea, and fungi detected in all 5,734 TCGA samples, which can serve as a public reference for future investigations.
- Research Article
22
- 10.1002/ijc.29512
- Apr 20, 2015
- International Journal of Cancer
microRNA regulation in cancer: One arm or two arms?
- Research Article
5
- 10.3390/cancers15194826
- Oct 1, 2023
- Cancers
Simple SummaryTumors are known to shed DNA into the bloodstream, and since the tumor DNA is marked by aberrant methylation patterns, this can be exploited for their detection through a simple blood sample. However, specific methylation biomarkers that efficiently detect a broad range of tumors and are effective at early-stage disease are still lacking. In this study we identify two novel methylation biomarkers and combine these with an already existing biomarker to improve multi-cancer detection. We test their performances as individual and combined markers using large methylation array datasets covering multiple cancer types, mimic blood samples using data from healthy blood cell DNA, and finally test the biomarkers in cancer plasma samples. We find that the combination of markers greatly improves the ability of the test to distinguish between cancer and normal samples, and in addition we provide the research field with a complete workflow for evaluating novel methylation biomarkers based on pre-existing datasets.The ability to detect several types of cancer using a non-invasive, blood-based test holds the potential to revolutionize oncology screening. We mined tumor methylation array data from the Cancer Genome Atlas (TCGA) covering 14 cancer types and identified two novel, broadly-occurring methylation markers at TLX1 and GALR1. To evaluate their performance as a generalized blood-based screening approach, along with our previously reported methylation biomarker, ZNF154, we rigorously assessed each marker individually or combined. Utilizing TCGA methylation data and applying logistic regression models within each individual cancer type, we found that the three-marker combination significantly increased the average area under the ROC curve (AUC) across the 14 tumor types compared to single markers (p = 1.158 × 10−10; Friedman test). Furthermore, we simulated dilutions of tumor DNA into healthy blood cell DNA and demonstrated increased AUC of combined markers across all dilution levels. Finally, we evaluated assay performance in bisulfite sequenced DNA from patient tumors and plasma, including early-stage samples. When combining all three markers, the assay correctly identified nine out of nine lung cancer plasma samples. In patient plasma from hepatocellular carcinoma, ZNF154 alone yielded the highest combined sensitivity and specificity values averaging 68% and 72%, whereas multiple markers could achieve higher sensitivity or specificity, but not both. Altogether, this study presents a comprehensive pipeline for the identification, testing, and validation of multi-cancer methylation biomarkers with a considerable potential for detecting a broad range of cancer types in patient blood samples.
- Research Article
103
- 10.1093/neuonc/not024
- Mar 15, 2013
- Neuro-Oncology
The Cancer Genome Atlas (TCGA) project is a large-scale effort with the goal of identifying novel molecular aberrations in glioblastoma (GBM). Here, we describe an in-depth analysis of gene expression data and copy number aberration (CNA) data to classify GBMs into prognostic groups to determine correlates of subtypes that may be biologically significant. To identify predictive survival models, we searched TCGA in 173 patients and identified 42 probe sets (P = .0005) that could be used to divide the tumor samples into 3 groups and showed a significantly (P = .0006) improved overall survival. Kaplan-Meier plots showed that the median survival of group 3 was markedly longer (127 weeks) than that of groups 1 and 2 (47 and 52 weeks, respectively). We then validated the 42 probe sets to stratify the patients according to survival in other public GBM gene expression datasets (eg, GSE4290 dataset). An overall analysis of the gene expression and copy number aberration using a multivariate Cox regression model showed that the 42 probe sets had a significant (P < .018) prognostic value independent of other variables. By integrating multidimensional genomic data from TCGA, we identified a specific survival model in a new prognostic group of GBM and suggest that molecular stratification of patients with GBM into homogeneous subgroups may provide opportunities for the development of new treatment modalities.