Single-cell analysis of mixed-lineage states leading to a binary cell fate choice.
Delineating hierarchical cellular states, including rare intermediates and the networks of regulatory genes that orchestrate cell-type specification, are continuing challenges for developmental biology. Single-cell RNA sequencing is greatly accelerating such research, given its power to provide comprehensive descriptions of genomic states and their presumptive regulators. Haematopoietic multipotential progenitor cells, as well as bipotential intermediates, manifest mixed-lineage patterns of gene expression at a single-cell level. Such mixed-lineage states may reflect the molecular priming of different developmental potentials by co-expressed alternative-lineage determinants, namely transcription factors. Although a bistable gene regulatory network has been proposed to regulate the specification of either neutrophils or macrophages, the nature of the transition states manifested in vivo, and the underlying dynamics of the cell-fate determinants, have remained elusive. Here we use single-cell RNA sequencing coupled with a new analytic tool, iterative clustering and guide-gene selection, and clonogenic assays to delineate hierarchical genomic and regulatory states that culminate in neutrophil or macrophage specification in mice. We show that this analysis captured prevalent mixed-lineage intermediates that manifested concurrent expression of haematopoietic stem cell/progenitor and myeloid progenitor cell genes. It also revealed rare metastable intermediates that had collapsed the haematopoietic stem cell/progenitor gene expression programme, instead expressing low levels of the myeloid determinants, Irf8 and Gfi1 (refs 9, 10, 11, 12, 13). Genetic perturbations and chromatin immunoprecipitation followed by sequencing revealed Irf8 and Gfi1 as key components of counteracting myeloid-gene-regulatory networks. Combined loss of these two determinants 'trapped' the metastable intermediate. We propose that mixed-lineage states are obligatory during cell-fate specification, manifest differing frequencies because of their dynamic instability and are dictated by counteracting gene-regulatory networks.
- Research Article
54
- 10.1016/j.exphem.2016.06.010
- Sep 1, 2016
- Experimental Hematology
Single-cell analysis of mixed-lineage states leading to a binary cell fate choice
- Supplementary Content
- 10.7907/nvwb-ap91.
- Jan 1, 2018
Development is an inherently dynamic process where cell fate specification occurs continuously and in a progressive manner. Thus, a major focus in developmental biology is solving the gene regulatory networks (GRNs) that underlie specification of cell fates. GRNs specify new spatial domains of cells by controlling the expression of their changing regulatory states throughout development. Regulatory states are composed of combinations of expressed regulatory genes which encode transcription factors that form regulatory circuits which function to carry out the specific developmental tasks involved in cell fate specification. To investigate the differences of GRNs operating in the embryo and their change over development, we sought to identify and characterize the regulatory states present in multiple developmental stages of sea urchin embryogenesis. We performed a genome-wide survey and embryo-wide annotation of regulatory gene expression by whole mount in situ hybridization at five consecutive developmental time-points in order to determine regulatory states and their developmental trajectory. We determined at least 74 distinct regulatory states expressed in discrete developmental domains which coincide with larval morphological structures and show that their progenitor domains foreshadow the ensuing larval morphology. Among these domains, we identified bilateral ciliary photoreceptors in the larva which express a distinct regulatory state that include factors known in ciliary photoreceptor specification. We show that this photoreceptor regulatory state does not express the genes of the retinal determination network that specify eyes in both flies and vertebrates. In addition, we show that though the sizes of regulatory states are comparable over developmental time, no two regulatory states are equal, even those expressed in a given domain at previous or subsequent developmental time-points. Lastly, we found that similarities among regulatory states reflect a common developmental function but not necessarily a common developmental history. The results suggest that the combinations of transcription factors defining regulatory states are both spatially and temporally dynamic in their progressive specification of cell fates during development and that regulatory state expression is tightly associated with the developing morphology of the larva.
- Peer Review Report
- 10.7554/elife.70416.sa1
- Jul 6, 2021
Decision letter: Single-cell RNA sequencing of the Strongylocentrotus purpuratus larva reveals the blueprint of major cell types and nervous system of a non-chordate deuterostome
- Research Article
51
- 10.7554/elife.51254.sa2
- Dec 16, 2019
- eLife
Understanding how gene expression programs are controlled requires identifying regulatory relationships between transcription factors and target genes. Gene regulatory networks are typically constructed from gene expression data acquired following genetic perturbation or environmental stimulus. Single-cell RNA sequencing (scRNAseq) captures the gene expression state of thousands of individual cells in a single experiment, offering advantages in combinatorial experimental design, large numbers of independent measurements, and accessing the interaction between the cell cycle and environmental responses that is hidden by population-level analysis of gene expression. To leverage these advantages, we developed a method for scRNAseq in budding yeast (Saccharomyces cerevisiae). We pooled diverse transcriptionally barcoded gene deletion mutants in 11 different environmental conditions and determined their expression state by sequencing 38,285 individual cells. We benchmarked a framework for learning gene regulatory networks from scRNAseq data that incorporates multitask learning and constructed a global gene regulatory network comprising 12,228 interactions.
- Research Article
177
- 10.7554/elife.51254
- Jan 27, 2020
- eLife
Understanding how gene expression programs are controlled requires identifying regulatory relationships between transcription factors and target genes. Gene regulatory networks are typically constructed from gene expression data acquired following genetic perturbation or environmental stimulus. Single-cell RNA sequencing (scRNAseq) captures the gene expression state of thousands of individual cells in a single experiment, offering advantages in combinatorial experimental design, large numbers of independent measurements, and accessing the interaction between the cell cycle and environmental responses that is hidden by population-level analysis of gene expression. To leverage these advantages, we developed a method for scRNAseq in budding yeast (Saccharomyces cerevisiae). We pooled diverse transcriptionally barcoded gene deletion mutants in 11 different environmental conditions and determined their expression state by sequencing 38,285 individual cells. We benchmarked a framework for learning gene regulatory networks from scRNAseq data that incorporates multitask learning and constructed a global gene regulatory network comprising 12,228 interactions.
- Research Article
10
- 10.1038/s41467-024-50822-y
- Aug 9, 2024
- Nature Communications
Cell fate specification occurs along invariant species-specific trajectories that define the animal body plan. This process is controlled by gene regulatory networks that regulate the expression of the limited set of transcription factors encoded in animal genomes. Here we globally assess the spatial expression of ~90% of expressed transcription factors during sea urchin development from embryo to larva to determine the activity of gene regulatory networks and their regulatory states during cell fate specification. We show that >200 embryonically expressed transcription factors together define >70 cell fates that recapitulate the morphological and functional organization of this organism. Most cell fate-specific regulatory states consist of ~15–40 transcription factors with similarity particularly among functionally related cell types regardless of developmental origin. Temporally, regulatory states change continuously during development, indicating that progressive changes in regulatory circuit activity determine cell fate specification. We conclude that the combinatorial expression of transcription factors provides molecular definitions that suffice for the unique specification of cell states in time and space during embryogenesis.
- Research Article
14
- 10.3389/fcell.2021.765578
- Nov 30, 2021
- Frontiers in Cell and Developmental Biology
Colorectal cancer (CRC) manifests as gastrointestinal tumors with high intratumoral heterogeneity. Recent studies have demonstrated that CRC may consist of tumor cells with different consensus molecular subtypes (CMS). The advancements in single-cell RNA sequencing have facilitated the development of gene regulatory networks to decode key regulators for specific cell types. Herein, we comprehensively analyzed the CMS of CRC patients by using single-cell RNA-sequencing data. CMS for all malignant cells were assigned using CMScaller. Gene set variation analysis showed pathway activity differences consistent with those reported in previous studies. Cell–cell communication analysis confirmed that CMS1 was more closely related to immune cells, and that monocytes and macrophages play dominant roles in the CRC tumor microenvironment. On the basis of the constructed gene regulation networks (GRNs) for each subtype, we identified that the critical transcription factor ERG is universally activated and upregulated in all CMS in comparison with normal cells, and that it performed diverse roles by regulating the expression of different downstream genes. In summary, molecular subtyping of single-cell RNA-sequencing data for colorectal cancer could elucidate the heterogeneity in gene regulatory networks and identify critical regulators of CRC.
- Research Article
2
- 10.1158/1538-7445.tim2013-a81
- Feb 1, 2013
- Cancer Research
Purpose: Invasion-metastasis cascades that underlie macroscopic metastases remain poorly understood. Tackling such a complexity demands systems biology approaches by means of filtering and integrating myriad data. To gain a fine resolution of mechanisms underlying the process of metastasis, here we reconstruct global gene regulatory and protein interaction networks that potentially drive breast cancer metastasis. Procedures: A systematic literature review and a comprehensive data mining of well-curated cancer gene datasets were carried out to uncover a core metastasis gene set (CMGS) using an in house Boolean logic framework. Meanwhile, a gene regulatory network (GRN) was built up to distinguish between key cancer driver genes and passenger genes through integration of multiple large-scale genomic studies of breast cancer. Furthermore, CMGS was projected onto the latest HRPD protein interaction network (PIN) and the GRN to derive metastasis-specific protein interaction network (mPIN) and gene regulatory network (mGRN), respectively. To gain mechanistic insights of mPIN and mGRN, a key network driver analysis was employed to identify critical hub nodes of the networks as key regulators. Results: We demonstrate that CMGS is of significantly higher connectivity in both PIN and GRN than those not in CMGS, suggesting its critical controlling power to the system. As expected, both mPIN and mGRN are most enriched for angiogenesis, integrin signaling, and p53 pathways. Moreover, these metastasis specific networks are also enriched for well-known metastasis-related pathways such as Wnt, TGF-beta, and FGF pathways. Surprisingly, inflammatory pathways, such as chemokine signaling and B-cell activation pathways are also over-represented in the networks. The top hub genes of mPIN include AR, ABL1, ESR1, AKT1, SMAD4, TP53, CSNK2A1, MAPK1, PIK3R1 and SMAD3. This hub gene set is distinct from the top key regulators of mGRN including ACTA2, ACTL6A, ADM, AEBP1, ARF1, ARPC4, ATM, AURKA and BIRC5. Notably, all the top ten drivers of mGRN except ATM and AURKA are not in CMGS, suggesting the novel targets of metastasis uncovered through our integrative network analysis. Indeed, ARF1, encoding the GTPase ADP-ribosylation factor 1, has been shown to play an important role in both cancer cell proliferation and migration. Moreover, the subnetwork regulated by ARF1 clearly links to inflammatory key players, especially via microvesicle related biological process. Importantly, both driver gene sets are tightly connected by TP53 and MAPK8 physically and genetically, suggesting the strong connection between p53 pathway and inflammatory response as MAPK8 activation is pivotal for chemokine mediated inflammation processes especially for TNF-alpha. Conclusion: Taken together, our global metastasis-specific protein interaction and gene regulation networks as well as their key regulators shed light on a potential trajectory of invasion-metastasis cascades that enable the progress of breast cancer metastasis, underscoring the instrumental role of a novel signaling network involving p53, microvesicle associated ARF1, and TNF-alpha mediated inflammatory response pathways. Citation Format: Bin Zhang, Yongzhong Zhao, Jun Zhu. Global gene regulatory and protein interaction networks in breast cancer metastasis. [abstract]. In: Proceedings of the AACR Special Conference on Tumor Invasion and Metastasis; Jan 20-23, 2013; San Diego, CA. Philadelphia (PA): AACR; Cancer Res 2013;73(3 Suppl):Abstract nr A81.
- Preprint Article
- 10.7287/peerj.preprints.3193v1
- Aug 26, 2017
Single cell studies increasing reveal myriad cellular subtypes beyond those postulated or observed through optical and fluorescence microscopy as well as DNA sequencing studies. While gene sequencing at the single cell level offer a path towards illuminating, in totality, the different subtypes of cells present, the technique nevertheless does not offer answers concerning the functional repertoire of the cell, which is defined by the collection of RNA transcribed from the genome. Known as the transcriptome, transcribed RNA defines the function of the cell as proteins or effector RNA molecules, while the genome is the collection of all information endowed in the cell type, expressed or not. Thus, a particular cell state, lineage, cell fate or cellular differentiation is more fully depicted by transcriptomic analysis compared to delineating the genomic context at the single cell level. While conceptually sound and could be analysed by contemporary single cell RNA sequencing technology and data analysis pipelines, the relative instability of RNA in view of RNase in the environment would make sample preparation particularly challenging, where degradation of cellular RNA by extraneous factors could provide a misinterpretation of specific functions available to a cell type. Hence, RNA as the de facto functional molecule of the cell defining the proteomics landscape as well as effector RNA repertoire, meant that RNA transcriptomics at the single cell level is the way forward if the goal is to understand all available cell types, lineage, cell fate and cellular differentiation. Given that a cell state is defined by the functions encoded by functional molecules such as proteins and RNA, single cell RNA sequencing offers a larger contextual basis for understanding cellular decision making and functions, for example, proteins are increasingly known to work in concert with RNA effector molecules in enabling a function. Hence, providing a view of the diverse cell types and lineages present in a body, single cell RNA sequencing is only hampered by the high sensitivity required to analyse the small amount of RNA available in single cells, as well as the perennial problem of RNA studies: how to prevent or reduce RNA degradation by environmental RNase enzymes. Ability to reduce RNA degradation would provide the cell biologist a unique view of the functional landscape of different cells in the body through the language of RNA.
- Book Chapter
7
- 10.1007/978-3-540-73871-8_26
- Jan 1, 2007
Bayesian networks are widely used to infer genes regulatory network from their transcriptional expression data. Bayesian network of the best score is usually chosen as genes regulatory model. However, without the hint from biological ground truth, and given a small number of transcriptional expression observations, the resulting Bayesian networks might not correspond to the real one. To deal with these two constrains, this paper proposes a stochastic approach to fit an existing hypothetical gene regulatory network, derived from biological evidence, with few available amount of transcriptional expression levels of the genes. The hypothetical gene regulatory network is set as an initial model of Bayesian network and fitted with transcriptional expression data by using Metropolis-Hastings algorithm. In this work, the transcriptional regulation of gene CYC1 by co-regulators HAP2 HAP3 HAP4 of yeast (Saccharomyces Cerevisiae) is considered as example. Due to the simulation results, ten probable gene regulatory networks which are similar to the given hypothetical model are obtained. This shows that Metropolis-Hastings algorithm can be used as a simulation model for gene regulatory network.
- Research Article
245
- 10.1016/j.stem.2020.01.013
- Feb 27, 2020
- Cell Stem Cell
Restraining Lysosomal Activity Preserves Hematopoietic Stem Cell Quiescence and Potency.
- Research Article
6
- 10.3760/cma.j.issn.1673-0860.2016.10.008
- Oct 7, 2016
- Zhonghua er bi yan hou tou jing wai ke za zhi = Chinese journal of otorhinolaryngology head and neck surgery
Objective: To analysis the important genes and functions of cochlear hair cells with oxidative stress injury, by the construction of gene regulatory network which based on different miRNA in cochlear hair cells in vitro with oxidative stress injury, and to explore the molecular mechanisms of deafness based on oxidative stress injury. Method: The oxidative stress damage cochlear hair cell model was induced by 200 μmol/L t-BHP exposure in vitro. Small RNA deep sequencing analyzed the difference expression of miRNA and contructed gene regulatory network by 6 most significant difference miRNA. The important interaction genes in regulatory network were screened and important genes function were annotated by GeneCards. Result: There were 24 different miRNAs in cochlear hair cells with oxidative stress injury by sRNA deep sequencing.Six most significant difference miRNA were: mir-1934 (logFC=2.367 947, P=2.35×10-7), mir-411 (logFC=2.093 687, P=3.13×10-6), mir-717 (logFC=1.927 67, P=3.24×10-5), mir-503 (logFC=-2.021 45, P=3.07×10-6), mir-467e (logFC=-1.953 28, P=0.000 137), and mir-699o (logFC=-1.950 06, P=0.000 517). Eleven important genes in miRNA regulatory network were: Akt1, Src, Ctnnb1, Creb1, Ccnd1, Egfr, Gsk3b, Pten, Cdh1, Fras1, and Ccnd2. Their main functions were to regulate hair cells apoptosis and proliferation by different intracellular signaling pathways. Conclusion: There are many signaling pathways (PI3K-AKT/PKB signaling pathway, AKT/PKB signaling pathway, Wnt signaling pathway, ERK signaling pathway, and Ras signaling pathway) involved in the regulation of apoptosis and proliferation in cochlear hair cells with oxidative stress injury and these signaling pathways are linked to each other to form a network. PI3K-AKT/PKB signaling pathway seems to be the most active in cochlear hair cells with oxidative stress injury.
- Preprint Article
- 10.7490/f1000research.1114795.1
- Aug 25, 2017
- F1000Research
Various methods for identifying differentially expressed genes yield quite different results. Here, we showed that key genes in regulatory networks derived by downstream analysis from lists of differentially expressed genes are a more robust alternative for the understanding of disease processes. As a basis for this, we used regulatory networks involving transcription factors, microRNAs, and target genes that we predicted with our TFmiR web server [1] from a set of differentially expressed genes and identified hotspot degree genes using this tool. Then we applied the ILP formulation of minimum dominating set [2] to find the key dominators in the underlying networks. Here, we processed RNA-Seq data taken from The Cancer Genome Atlas (TCGA) for matched tumor and normal samples of hepatocellular carcinoma (liver cancer) and breast cancer patients. Differentially expressed genes were identified by four different bioinformatics tools. While the overlap between the sets of significant differentially expressed genes was only 26 % in liver cancer and 28 % in breast cancer, we found that the topology of the regulatory networks constructed using TFmiR for the different sets of differentially expressed genes was highly similar with respect to hub degree nodes and dominators. This suggests that key genes identified in regulatory networks derived from differentially expressed genes may be a more robust basis for understanding diseases processes than simply inspecting the lists of differentially expressed genes. [1] TFmiR: A Web Server for Constructing and Analyzing Disease-specific Transcription Factor and miRNA co-regulatory Networks. Mohamed Hamd, Christian Spaniol, Maryam Nazarieh , Volkhard Helms . Nucleic Acid Research, 2015. Available at: http://service.bioinformatik.uni-saarland.de/tfmir [2] Identification of Key Player Genes in Gene Regulatory Networks. Maryam Nazarieh , Andreas Wiese, Thorsten Will, Mohamed Hamed, Volkhard Helms . BMC System Biology, 2016. Available at: http://apps.cytoscape.org/apps/mcds , https://github.com/maryamNazarieh/KeyRegulatoryGenes
- Research Article
- 10.34133/csbj.0080
- Jan 1, 2026
- Computational and structural biotechnology journal
TDAGENE: Inference of Gene Regulatory Network Based on Topological Data Analysis and Graph Attention Network for Single-Cell RNA Sequencing Data.
- Research Article
66
- 10.1093/bib/bbad414
- Sep 22, 2023
- Briefings in Bioinformatics
Single-cell RNA-sequencing (scRNA-seq) has emerged as a powerful technique for studying gene expression patterns at the single-cell level. Inferring gene regulatory networks (GRNs) from scRNA-seq data provides insight into cellular phenotypes from the genomic level. However, the high sparsity, noise and dropout events inherent in scRNA-seq data present challenges for GRN inference. In recent years, the dramatic increase in data on experimentally validated transcription factors binding to DNA has made it possible to infer GRNs by supervised methods. In this study, we address the problem of GRN inference by framing it as a graph link prediction task. In this paper, we propose a novel framework called GNNLink, which leverages known GRNs to deduce the potential regulatory interdependencies between genes. First, we preprocess the raw scRNA-seq data. Then, we introduce a graph convolutional network-based interaction graph encoder to effectively refine gene features by capturing interdependencies between nodes in the network. Finally, the inference of GRN is obtained by performing matrix completion operation on node features. The features obtained from model training can be applied to downstream tasks such as measuring similarity and inferring causality between gene pairs. To evaluate the performance of GNNLink, we compare it with six existing GRN reconstruction methods using seven scRNA-seq datasets. These datasets encompass diverse ground truth networks, including functional interaction networks, Loss of Function/Gain of Function data, non-specific ChIP-seq data and cell-type-specific ChIP-seq data. Our experimental results demonstrate that GNNLink achieves comparable or superior performance across these datasets, showcasing its robustness and accuracy. Furthermore, we observe consistent performance across datasets of varying scales. For reproducibility, we provide the data and source code of GNNLink on our GitHub repository: https://github.com/sdesignates/GNNLink.