AACR Project GENIE: Powering Precision Medicine through an International Consortium.
The AACR Project GENIE is an international data-sharing consortium focused on generating an evidence base for precision cancer medicine by integrating clinical-grade cancer genomic data with clinical outcome data for tens of thousands of cancer patients treated at multiple institutions worldwide. In conjunction with the first public data release from approximately 19,000 samples, we describe the goals, structure, and data standards of the consortium and report conclusions from high-level analysis of the initial phase of genomic data. We also provide examples of the clinical utility of GENIE data, such as an estimate of clinical actionability across multiple cancer types (>30%) and prediction of accrual rates to the NCI-MATCH trial that accurately reflect recently reported actual match rates. The GENIE database is expected to grow to >100,000 samples within 5 years and should serve as a powerful tool for precision cancer medicine.Significance: The AACR Project GENIE aims to catalyze sharing of integrated genomic and clinical datasets across multiple institutions worldwide, and thereby enable precision cancer medicine research, including the identification of novel therapeutic targets, design of biomarker-driven clinical trials, and identification of genomic determinants of response to therapy. Cancer Discov; 7(8); 818-31. ©2017 AACR.See related commentary by Litchfield et al., p. 796This article is highlighted in the In This Issue feature, p. 783.
- Preprint Article
- 10.1158/2159-8290.c.6546577.v1
- Apr 3, 2023
<div>Abstract<p>The AACR Project GENIE is an international data-sharing consortium focused on generating an evidence base for precision cancer medicine by integrating clinical-grade cancer genomic data with clinical outcome data for tens of thousands of cancer patients treated at multiple institutions worldwide. In conjunction with the first public data release from approximately 19,000 samples, we describe the goals, structure, and data standards of the consortium and report conclusions from high-level analysis of the initial phase of genomic data. We also provide examples of the clinical utility of GENIE data, such as an estimate of clinical actionability across multiple cancer types (>30%) and prediction of accrual rates to the NCI-MATCH trial that accurately reflect recently reported actual match rates. The GENIE database is expected to grow to >100,000 samples within 5 years and should serve as a powerful tool for precision cancer medicine.</p><p><b>Significance:</b> The AACR Project GENIE aims to catalyze sharing of integrated genomic and clinical datasets across multiple institutions worldwide, and thereby enable precision cancer medicine research, including the identification of novel therapeutic targets, design of biomarker-driven clinical trials, and identification of genomic determinants of response to therapy. <i>Cancer Discov; 7(8); 818–31. ©2017 AACR.</i></p><p><i>See related commentary by Litchfield et al., p. 796</i>.</p><p><i>This article is highlighted in the In This Issue feature, p. 783</i></p></div>
- Preprint Article
- 10.1158/2159-8290.c.6546577
- Apr 3, 2023
<div>Abstract<p>The AACR Project GENIE is an international data-sharing consortium focused on generating an evidence base for precision cancer medicine by integrating clinical-grade cancer genomic data with clinical outcome data for tens of thousands of cancer patients treated at multiple institutions worldwide. In conjunction with the first public data release from approximately 19,000 samples, we describe the goals, structure, and data standards of the consortium and report conclusions from high-level analysis of the initial phase of genomic data. We also provide examples of the clinical utility of GENIE data, such as an estimate of clinical actionability across multiple cancer types (>30%) and prediction of accrual rates to the NCI-MATCH trial that accurately reflect recently reported actual match rates. The GENIE database is expected to grow to >100,000 samples within 5 years and should serve as a powerful tool for precision cancer medicine.</p><p><b>Significance:</b> The AACR Project GENIE aims to catalyze sharing of integrated genomic and clinical datasets across multiple institutions worldwide, and thereby enable precision cancer medicine research, including the identification of novel therapeutic targets, design of biomarker-driven clinical trials, and identification of genomic determinants of response to therapy. <i>Cancer Discov; 7(8); 818–31. ©2017 AACR.</i></p><p><i>See related commentary by Litchfield et al., p. 796</i>.</p><p><i>This article is highlighted in the In This Issue feature, p. 783</i></p></div>
- Research Article
1
- 10.1158/1538-7445.am2017-lb-102
- Jul 1, 2017
- Cancer Research
AACR Project Genomics Evidence Neoplasia Information Exchange (GENIE) is a multi-phase, multi-year, international data-sharing consortium whose goal is to generate an evidence base for precision cancer medicine by integrating and linking clinical-grade cancer genomic data with clinical outcome data for tens of thousands of cancer patients treated at multiple institutions worldwide. The project fulfills an unmet need in oncology by providing the statistical power necessary to identify novel therapeutic targets, to understand genomic determinants of response to therapy, to design new biomarker-driven clinical trials and ultimately, to improve clinical decision-making and the care delivered to patients. Here we describe the goals, structure and data standards of the GENIE consortium and conclusions from a high-level analysis of the first public release of genomic and limited clinical data from approximately 19,000 patients treated at eight cancer centers obtained during this initial phase of the project. We also explore the clinical utility of these genomic data by examining rates of clinical actionability across multiple cancer types and by estimating patient enrollment rates to the NCI MATCH Trial. Based on yearly rates of sequencing at each of the eight founding institutions, together with the planned addition of new members, we estimate the GENIE database could grow to &gt;100,000 samples within five years. Consistent with the goals of the proposed Cancer Moonshot National Cancer Data Ecosystem, GENIE is committed to the principles of generating interoperable, open access data that can be widely shared across the entire scientific community. Citation Format: Ethan Cerami, Alexander S. Baras, Justin Guinney, Eva Lepisto, Trevor J. Pugh, Nikolaus Schultz, Thomas Stricker, Shawn M. Sweeney, Laura J. van't Veer, Gerrit A. Meijer, Fabrice Andre, Victor E. Velculescu, Kenna R. Shaw, Mia A. Levy, Philippe L. Bedard, Barrett J. Rollins, Charles L. Sawyers, on behalf of the AACR Project GENIE Consortium. Landscape analysis of the initial data release from AACR Project GENIE [abstract]. In: Proceedings of the American Association for Cancer Research Annual Meeting 2017; 2017 Apr 1-5; Washington, DC. Philadelphia (PA): AACR; Cancer Res 2017;77(13 Suppl):Abstract nr LB-102. doi:10.1158/1538-7445.AM2017-LB-102
- Abstract
- 10.1136/jitc-2022-sitc2022.1273
- Nov 1, 2022
- Journal for ImmunoTherapy of Cancer
<h3>Background</h3> Despite the clinical success of immune checkpoint blockade therapies, many patients do not respond to treatments or become resistant. Previous attempts to predict treatment efficacy suffered from limited accuracy...
- Research Article
- 10.1158/1538-7445.am2020-2058
- Aug 13, 2020
- Cancer Research
My Cancer Genome (https://www.mycancergenome.org; MCG) recently released a major update to the website, with an ultimate goal of adding of AACR Project GENIE data to demonstrate frequency of genetic variants in cancer, estimate matching to clinical trial arms, and augment the case report database. This abstract presents use of GENIE data 1) as part of a new process to automatically select cancer types and biomarkers for inclusion and 2) to display prevalence data on the website. My Cancer Genome is a publicly available precision cancer medicine knowledgebase managed by the Vanderbilt-Ingram Cancer Center. The mission of MCG is to curate and disseminate knowledge regarding the clinical significance of genomic alterations in cancer. The website was first made available to the public in early 2011 and has since grown into a resource visited more than 7,000 times a week by people from well over 200 countries and territories around the world. The new website was released in June 2019, and it includes several new features: a new clinical trials search using manually curated disease and biomarker eligibility criteria, therapeutic assertions illustrating situations where having or not having a particular biomarker detected supports or does not support use of a particular drug, and the inclusion of data from AACR Project GENIE. Addition of GENIE data to MCG was approved by AACR Project GENIE's Data Use and Publications Committee. First, any disease or biomarker associated with a therapeutic assertion or clinical trial has a page on MCG. In addition, any disease or biomarker that appears five times or more in the GENIE dataset is also given a page. In this way, the website will remain up-to-date with almost all diseases and biomarkers that can be considered relevant as precision oncology knowledge and GENIE grow. Next, GENIE data appears on disease and biomarker pages, in both charts and descriptive text. On disease pages, such as the breast carcinoma page, GENIE data appear in a chart of most commonly altered genes in breast carcinoma and most common alterations in breast carcinoma; in addition, text describing the prevalence of mutations appears under each gene listed in the section on significant genes in breast carcinoma. On biomarker pages, such as the ABL1 page, GENIE data appear in a chart of the most common diseases where ABL1 has been found to be altered and the most common alterations in ABL1 that appear in GENIE cases; in addition, text describing the prevalence of ABL1 mutations appears under each disease listed in the section on the significance of ABL1 in diseases. The addition of data from AACR Project GENIE to My Cancer Genome has improved and enhanced MCG's content, ability to scale, and ability to maintain. Finally, inclusion of data from AACR Project GENIE on My Cancer Genome provides a simple way to select and view subsets of GENIE data for specific diseases and biomarkers in the context of relevant precision oncology therapies and clinical trials. Citation Format: Christine M. Micheel, Michele L. LeNoue-Newton, Marilyn E. Holt, Neha Jain, Mia A. Levy. Augmenting the My Cancer Genome knowledge resource with data from AACR Project GENIE [abstract]. In: Proceedings of the Annual Meeting of the American Association for Cancer Research 2020; 2020 Apr 27-28 and Jun 22-24. Philadelphia (PA): AACR; Cancer Res 2020;80(16 Suppl):Abstract nr 2058.
- Abstract
- 10.1016/j.annonc.2022.07.1436
- Sep 1, 2022
- Annals of Oncology
1304P Characterizing the clinico-genomic landscape and outcomes of KRAS G12C mutated pancreas cancer
- Research Article
- 10.1158/1538-7445.am2023-lb056
- Apr 14, 2023
- Cancer Research
A key challenge in rare tumor research is the paucity of genomic data that can be used to understand and devise better therapeutic strategies for rare cancers. Furthermore, the “curse of dimensionality,” in which data has many features, such as genetic variants, but few specimens, makes it difficult or impossible to use conventional machine learning techniques to explore these data. To address these challenges in the context of tumors associated with the rare disease neurofibromatosis, we ran a hackathon to stimulate the development of methods to better understand the biology of tumors related to this disease. The hackathon had three challenges centered around variant effect prediction, drug discovery, and genomics. The genomics track of the hackathon leveraged the AACR Project GENIE database (1) and challenged participants to develop new frameworks that accurately use GENIE data to classify neurofibromatosis-related tumors. They were asked to first identify the neurofibromatosis-related tumors in the dataset. They were then asked to use one or more novel classification methods to classify the tumor samples into different groups based on genetic features. To help them do this, we provided access to version 13 of the GENIE database to the hackathon participants, though they were allowed to integrate other relevant datasets. The expected output was a classification method that differentiates different types of NF1, NF2, and schwannomatosis-related tumors using clinical sequencing data, as well as a list of the most important features in the algorithm for differentiating tumor types. Domain expert judges qualitatively scored each team’s rationale for defining and including “NF-related tumors” in their project, and scored the feature list based on the presence of known important biomarkers and features in NF tumors as well as potentially novel features that the algorithm identified. A technical judge also scored the code repository based on documentation and clarity of code. Two teams from the GENIE subchallenge were awarded prizes - Team Next GeNLP as the best overall GENIE challenge submission, and team “Artificial Intelligence for neurofibromatosis” for best project documentation. Both winning teams used methods based on natural language processing (NLP) techniques to reduce the dimensionality and complexity of the variant data, and to identify new representations of NF-relevant tumors, and then applied downstream analysis methods such as distance calculations and feature prioritization to better understand the genomic profiles of different tumors. While these methods and tools focused on NF-specific tumor types, we anticipate that they could be re-used by others to better explore the biology and interrelatedness of other rare tumors within the GENIE database. (1) The AACR Project GENIE Consortium. AACR Project GENIE: Powering Precision Medicine Through An International Consortium, Cancer Discov. 2017. The authors would like to acknowledge the American Association for Cancer Research and its financial and material support in the development of the AACR Project GENIE registry, as well as members of the consortium for their commitment to data sharing. Interpretations are the responsibility of study authors. Citation Format: Robert J. Allaway, Sasha Scott, Gabriel Altay, Hariprasad Donthi, Karthika R, Muhammad Alaa Alwattar, Lucas Pastur Romay, Ayesha Parsha, Jineta Banerjee, Chelsea Nayan, Julie Bletz, Salvatore La Rosa. Crowdsourcing rare cancer research in the Hack4NF GENIE-NF tumor identification and classification challenge [abstract]. In: Proceedings of the American Association for Cancer Research Annual Meeting 2023; Part 2 (Clinical Trials and Late-Breaking Research); 2023 Apr 14-19; Orlando, FL. Philadelphia (PA): AACR; Cancer Res 2023;83(8_Suppl):Abstract nr LB056.
- Research Article
- 10.1158/1538-7445.am2024-3519
- Mar 22, 2024
- Cancer Research
Background: Molecularly and clinically well-annotated patient datasets are ideal for studying tumor biology and developing robust machine learning (ML) models for predicting outcome and treatment response. These data however rarely exist in real-world settings or in sufficient quantities within research contexts. Large publicly available datasets like The Cancer Genome Atlas (TCGA) which provid multi-omic profiles for diverse cancer types, have profoundly advanced cancer research and facilitated development of novel therapies and personalized medicines. However, the absence of patient outcome data tied to treatment limits the applicability of these data for understanding and modeling treatment response. Real-world clinicogenomics cohorts, such as the AACR Project GENIE, on the other hand are typically very rich in clinical annotations, including treatment regimens and outcomes measures. These data are however sparsely annotated for patient tumor molecular profiles, rarely exceeding ~100’s of genes profiled. We hypothesized that it would be possible to reconstruct latent tumor mRNA representations from limited genomic and clinical data available in real-world clinicogenomic cohorts, and that these reconstructed expression profiles would be useful for a variety of clinically meaningful downstream applications. Methods: We developed an ML model, called Mut2Ex, to reconstruct tumor gene expression profiles using genetic information available on commercial next generation sequencing panels using a Principle Label Space Transformation (PLST) we adapted to regression problem, along with embeddings from clinical information (OncoTree code, sex and stage) generated by a language model. Mut2Ex was trained on ~1200 cell lines from DepMap representing 26 cancer types, to generate ~2000 reconstructed mRNA gene profiles that were applied to a variety of clinical tasks. We used Mut2Ex to reconstruct mRNA profiles for ~10,000 tumors from TCGA and ~180,000 tumors from AACR Project GENIE. Results: Reconstructed mRNA expression by Mut2Ex was highly correlated with true expression in cell lines (rho = 0.926, [0.924-0.928 95% CI, N = 1184]). Compared to true expression, reconstructed profiles recapitulate sub-clusters within cancer types, PAM50 subtyping in breast tumors, survival signatures in colorectal tumors and multiple oncogenic signatures in a pan-cancer manner. Analysis of reconstructed expression for AACR Project GENIE tumors revealed expected enrichment of known driver genes within expression subtypes and enrichment of oncogenic signatures associated with distinct clinical outcomes in a cancer type specific manner. Conclusions: Our flexible analytic framework for reconstructing gene expression profiles from clinicogenomics data substantially augments the clinical utility and value of data acquired in real-world settings. Citation Format: Maayan Baron, Sunil Kumar, Felicia Kuperwaser, Dillon Tracy, Emily Vucic, Jeff Sherman. Reconstructing a latent representation of gene expression from genomic alterations to improve clinical utility of real-world clinicogenomics data [abstract]. In: Proceedings of the American Association for Cancer Research Annual Meeting 2024; Part 1 (Regular Abstracts); 2024 Apr 5-10; San Diego, CA. Philadelphia (PA): AACR; Cancer Res 2024;84(6_Suppl):Abstract nr 3519.
- Research Article
- 10.1158/1538-7445.genfunc25-a047
- Mar 11, 2025
- Cancer Research
Access to well-characterized patient-derived cancer models (PDCMs) is crucial for advancing our understanding of cancer biology and developing targeted therapies. CancerModels.Org addresses this need by providing a unified, open global research platform for PDCMs. Its key objectives are to aggregate and standardize clinical, genomic, and functional data from various patient-derived xenografts (PDXs), organoids, and cell lines across multiple cancer types and to facilitate easier discovery and comparison of cancer models by implementing standardized data formats and adhering to FAIR (Findable, Accessible, Interoperable, and Reusable) data principles. CancerModels.Org, launched in 2023, offers a collection of over 9300 PDCMs across 13 cancer types and includes rare pediatric PDX models and models from minority ethnic backgrounds. A total of over 365 billion data points from 55 model providers in 12 countries are made available across a variety of data types, including gene expression, gene mutation, copy number alteration, biomarkers, immune-related features of tumors, patient treatment, model treatment and imaging data (Release 6.6 - Oct 2024), making CancerModels.Org the largest free-to-consumer and open-access resource of this kind. The platform's user-friendly web interface allows for efficient searching and filtering of models based on various criteria, including cancer type, genomic alterations, patient clinical information, and treatment response. The data is also accessible via cBioPortal instance and REST API, enabling offline analyses. For an improved prioritization of PDCMs we performed knowledge enrichment by linking to external resources, such as publication platforms, cancer-specific annotation tools (COSMIC, CIViC, OncoMX, OpenCRAVAT, ClinGen), and raw data archives (ENA, EGA, GEO, dbGAP - for sequence, and BioImage archive for image data).The project is driving the development of and promoting the use of descriptive standards to facilitate data interoperability and promote global sharing of models, e.g. PDX Minimal Information standard (PDX-MI) and in vitro PDCM Minimal Information standard (manuscript in preparation). We provide expertise and software components to support several worldwide consortia, including PDXNet and EurOPDX. The project is supported by NCI and is freely available on Github under an Apache 2.0 licence.In conclusion, CancerModels.Org serves as a valuable resource for researchers maximizing utility and reusability of models and data, facilitating collaboration and data sharing among researchers and fostering a more integrated approach to precision oncology. By providing a standardized format for model data and encouraging contributions from the scientific community, the platform promotes reproducibility and accelerates the translation of laboratory findings into clinical applications, contributing to the development of more effective, personalized cancer treatments. Citation Format: Zinaida Perova, Mauricio Martinez, Tushar Mandloi, Marcelo Rios Almanza, Steven Neuhauser, Debbie Krupke, Dale Begley, Carol Bult, Helen Parkinson. CancerModels.Org: A comprehensive resource for advancing precision medicine in cancer research [abstract]. In: Proceedings of the AACR Special Conference in Cancer Research: Functional and Genomic Precision Medicine in Cancer: Different Perspectives, Common Goals; 2025 Mar 11-13; Boston, MA. Philadelphia (PA): AACR; Cancer Res 2025;85(5 Suppl):Abstract nr A047.
- Research Article
106
- 10.1002/1878-0261.12465
- Feb 22, 2019
- Molecular Oncology
Cancer treatment has made significant strides towards the promise of personalized medicine. Recent scientific advances have shown that there are numerous genetic deregulations that are common in multiple cancer types, raising the possibility of developing drugs targeting those deregulations irrespective of the tumour type. Precision Cancer Medicine (PCM) was born out of accumulated evidence matching targeted agents with these tumour molecular deregulations. At the same time, the therapeutic armamentarium is rapidly increasing and the number of new drugs (including immune‐oncology agents) entering drug development continues to rise. These factors, added to strong collaboration with regulatory agencies, which have approved novel agents based on data obtained from phase 1/2 trials, have led to unprecedented evolution in the design of early‐stage clinical trials. Currently, we have seen rapid phase 1 dose‐escalation trials followed by remarkably large expansion cohorts, and are witnessing the emergence of new trials, such as adaptive studies with basket and umbrella designs aimed at optimizing the biomarker–drug co‐development process. Alongside the growing complexity of these clinical trials, new frameworks for stronger and faster collaboration between all stakeholders in drug development, including academic institutions and frameworks, clinicians, pharma companies and regulatory agencies, have been established. In this review article, we describe the main challenges and opportunities that these new trial designs may provide for a more efficient drug development process, which may ultimately help ensure that PCM becomes a reality for patients.
- Research Article
1
- 10.2302/kjm.66-003-abst
- Jan 1, 2017
- The Keio journal of medicine
Precision Cancer Medicine: Tumors are classified into several molecular subtypes by Genomic Sequencing. Comprehensive targeted-gene panel provides more therapeutic options. CANCERPLEX-JP version 4.0 includes 435 actionable genes, which clinical approaches are available. Utilizing the gene panel platform, we assessed the genes and pathways most frequently altered in Japanese and US cases (Genome Med 2016;8:136). We demonstrate concordance of CANCERPLEX-JP with whole-exome sequencing from the TCGA in identifying hypermutated samples and microsatellite instability in multiple cancer types, such as colorectal and gastric cancer. We introduced our activity for precision medicine at 4th US-Japan Clinical Trials in Oncology Workshop sponsored by Embassy of Japan, at Washington DC, held in June 9, 2016. We highlight the clinical utility of CANCERPLEX-JP (435 genes) in guiding treatment strategies with targeted therapy in solid tumors, thus providing rationale for Comprehensive Genomic Sequencing in actualizing precision medicine. The goal of Precision Medicine is cost effectiveness and realization of genome drug discovery.Super-computing System: This introduction plan is based on the massive medical big data (cancer gene mutation information, genomic data, image data, etc.) possessed by Niigata Prefecture and Niigata University, to construct an integrated analysis system incorporating deep learning, and introduces a high-performance computer system for the development and activation of related industries. Therefore, it is necessary to have an operation system that can connect to the hospital electronic medical record without interference.(Presented at the 1943rd Meeting, July 5, 2017).
- Research Article
- 10.1158/1538-7445.am2023-4265
- Apr 4, 2023
- Cancer Research
Introduction: IMM-1-104, with pan-RAS activity through deep cyclic inhibition MEK, was evaluated in humanized 3D preclinical tumor models displaying diverse MAPK pathway activation events. Based on drug-response, sensitivity and resistance profiles, a biomarker signature for IMM-1-104 was developed in order to project potential therapeutic response of cancer patients found in the AACR Project GENIE (GENIE) database. Experimental Procedures: Humanized 3D preclinical models better predict in vivo tumor responses versus 2D culture and more accurately replicate biology of human tumors. Therefore, the antitumor activity of IMM-1-104 was evaluated in over 130 tumor models spanning 12 distinct histologies in the humanized 3D tumor growth assay (3D-TGA). Cell-based whole exome sequencing readouts were combined with 3D-TGA results to build a pharmacogenomic response algorithm. When applied to the GENIE patient database, resultant tumor-specific response landscapes helped to inform an early pan-RAS clinical trial design for IMM-1-104. Summary of New Data: A machine learning model was developed to predict IMM-1-104 sensitivity using response-associated genes and signaling networks that were identified using 3D-TGA pharmacogenomics data. This model was used to estimate GENIE patient IMM-1-104 response profiles across key solid tumor indications. In addition, mutation constellations from GENIE were compared with those observed in cell lines to identify preclinical models that best resemble real-world patients. This effort was designed to further enrich the translational fidelity of specific tumor models with the goal of translationally identifying patient populations most likely to benefit from IMM-1-104 treatment. Conclusions: The depth of response to IMM-1-104 was evaluated across a panel of diverse 3D-TGA tumor models and led to identification of a biomarker signature for therapeutically addressable MAPK pathway addiction. To translate these findings into a relevant clinical application, a response algorithm was developed and applied to the GENIE database, which has cataloged the molecular profiles of over 100,000 cancer patients. Mutational landscapes of patients within GENIE helped identify preclinical models that better represent patient profiles likely to be encountered in the clinic. This approach could, as a general principle, be applied as a tool for improving biomarker discovery and clinical translation of oncology drugs. Citation Format: Praveen Nair, Sarah Kolitz, Jason Funt, Peter J. King, Kevin D. Fowler, Anna Travesa, Ian Rose, John Brothers, Amy Axel, Scott Barrett, Benjamin J. Zeskind, Brett M. Hall. Humanized 3D tumor models that are mutationally aligned with AACR GENIE patients predict IMM-1-104 activity in RAS-addicted tumors. [abstract]. In: Proceedings of the American Association for Cancer Research Annual Meeting 2023; Part 1 (Regular and Invited Abstracts); 2023 Apr 14-19; Orlando, FL. Philadelphia (PA): AACR; Cancer Res 2023;83(7_Suppl):Abstract nr 4265.
- Preprint Article
- 10.1158/2159-8290.22531885
- Apr 3, 2023
<p>Table S4</p>
- Preprint Article
- 10.1158/2159-8290.22531897.v1
- Apr 3, 2023
<p>AACR GENIE Data Guide</p>
- Preprint Article
- 10.1158/2159-8290.22531888
- Apr 3, 2023
<p>Table S3</p>