Cytoscape 2.8: new features for data integration and network visualization
Summary: Cytoscape is a popular bioinformatics package for biological network visualization and data integration. Version 2.8 introduces two powerful new features—Custom Node Graphics and Attribute Equations—which can be used jointly to greatly enhance Cytoscape's data integration and visualization capabilities. Custom Node Graphics allow an image to be projected onto a node, including images generated dynamically or at remote locations. Attribute Equations provide Cytoscape with spreadsheet-like functionality in which the value of an attribute is computed dynamically as a function of other attributes and network properties.Availability and implementation: Cytoscape is a desktop Java application released under the Library Gnu Public License (LGPL). Binary install bundles and source code for Cytoscape 2.8 are available for download from http://cytoscape.org.Contact: msmoot@ucsd.edu
- Research Article
6
- 10.1093/bioinformatics/btt377
- Jul 26, 2013
- Bioinformatics
Network-level visualization of functional data is a key aspect of both analysis and understanding of biological systems. In a continuing effort to create clear and integrated visualizations that facilitate the gathering of novel biological insights despite the overwhelming complexity of data, we present here the GrAph LANdscape VisualizaTion (GALANT), a Cytoscape plugin that builds functional landscapes onto biological networks. By using GALANT, it is possible to project any type of numerical data onto a network to create a smoothed data map resembling the network layout. As a Cytoscape plugin, GALANT is further improved by the functionalities of Cytoscape, the popular bioinformatics package for biological network visualization and data integration. http://www.lbbc.ibb.unesp.br/galant.
- Research Article
13
- 10.14719/pst.2014.1.2.26
- Apr 1, 2014
- Plant Science Today
Knowledge of the interactome improves the understanding of disease metabolism.Biological information about interactions among genes and their protein products, computationally extracted in the context of SysBiomics, can hint at molecular causes of diseases, be essential for understanding biological systems, and provide clues for new therapeutic approaches.Quick and efficient access to this data have become critical issues for biologists.We have implemented a computational platform that integrates pathway, protein-protein interaction, differentially expressed genome and literature mining data to result in comprehensive networks for insomnia and intervention effects of Jujuboside B (JuB).The interaction data were imported into Cytoscape software, a popular bioinformatics package for biological network visualization and data integration, for screening the central nodes of the network, exploiting functional study of the central node genes, exploring the mechanism of insomnia.Results showed that seven differentially expressed genes confirmed by Cytoscape as the central nodes of the network in insomnia had interactions, forming a complicated interaction network (77 nodes, 96 edges).Among gene nodes, HBA1, LEP, MAOA, PRNP, GHRL, CLOCK and SLC6A4 were verified as the genes with maximal differential expressions.Of note, we further observed that the HBA1, LEP, SLC6A4 and MAOA were JuB target genes.The interaction network of the differentially expressed genes, especially the central nodes of this
- Book Chapter
9
- 10.1007/3-540-45581-7_20
- Jan 1, 2001
Evaluation ofd ata non-quality in database or datawarehouse systems is a preliminary stage before any data usage and analysis, moreover in the context ofd ata integration where several sources provide more or less redundant or contradictory information items and whose quality is often unknown, imprecise and very heterogeneous. Our application domain is bioinformatics where more than five hundred of semi-structured databanks propose biological information without any quality information (i.e. metadata and statistics describing the production and the management oft he biological data). In order to facilitate the multisource data integration in the context ofd istributed biological databanks, we propose a technique based on the concepts ofq uality contract and data source negotiation for a standard wrapper-mediator architecture. A quality source contract allows to specify quality dimensions necessary to the mediator for data extraction among several distributed resources. The source selection is dynamically computed with the contract negotiation which we propose to include into the mediation and the global query processings before data acquisition. The integration of the multisource biological data is differed for the restitution and combination of the results ofthe global user's query by techniques ofda ta recommendation taking into account source quality requirements.
- Book Chapter
1
- 10.4018/978-1-60960-067-9.ch014
- Jan 1, 2011
The volume of information derived from post genomic technologies is rapidly increasing. Due to the amount of involved data, novel computational methods are needed for the analysis and knowledge discovery into the massive data sets produced by these new technologies. Furthermore, data integration is also gaining attention for merging signals from different sources in order to discover unknown relations. This chapter presents a pipeline for biological data integration and discovery of a priori unknown relationships between gene expressions and metabolite accumulations. In this pipeline, two standard clustering methods are compared against a novel neural network approach. The neural model provides a simple visualization interface for identification of coordinated patterns variations, independently of the number of produced clusters. Several quality measurements have been defined for the evaluation of the clustering results obtained on a case study involving transcriptomic and metabolomic profiles from tomato fruits. Moreover, a method is proposed for the evaluation of the biological significance of the clusters found. The neural model has shown a high performance in most of the quality measures, with internal coherence in all the identified clusters and better visualization capabilities.
- Book Chapter
- 10.4018/978-1-4666-2455-9.ch011
- Jan 1, 2013
The volume of information derived from post genomic technologies is rapidly increasing. Due to the amount of involved data, novel computational methods are needed for the analysis and knowledge discovery into the massive data sets produced by these new technologies. Furthermore, data integration is also gaining attention for merging signals from different sources in order to discover unknown relations. This chapter presents a pipeline for biological data integration and discovery of a priori unknown relationships between gene expressions and metabolite accumulations. In this pipeline, two standard clustering methods are compared against a novel neural network approach. The neural model provides a simple visualization interface for identification of coordinated patterns variations, independently of the number of produced clusters. Several quality measurements have been defined for the evaluation of the clustering results obtained on a case study involving transcriptomic and metabolomic profiles from tomato fruits. Moreover, a method is proposed for the evaluation of the biological significance of the clusters found. The neural model has shown a high performance in most of the quality measures, with internal coherence in all the identified clusters and better visualization capabilities.
- Research Article
57
- 10.1186/1471-2105-7-286
- Jun 6, 2006
- BMC Bioinformatics
BackgroundThe biological information in genomic expression data can be understood, and computationally extracted, in the context of systems of interacting molecules. The automation of this information extraction requires high throughput management and analysis of genomic expression data, and integration of these data with other data types.ResultsSBEAMS-Microarray, a module of the open-source Systems Biology Experiment Analysis Management System (SBEAMS), enables MIAME-compliant storage, management, analysis, and integration of high-throughput genomic expression data. It is interoperable with the Cytoscape network integration, visualization, analysis, and modeling software platform.ConclusionSBEAMS-Microarray provides end-to-end support for genomic expression analyses for network-based systems biology research.
- Research Article
5
- 10.1177/1177932220906168
- Jan 1, 2020
- Bioinformatics and Biology Insights
Nowadays, the integration of biological data is a major challenge for bioinformatics. Many studies have examined gene expression in the epithelial tissue in the intestines of infants born to term and breastfed, generating a large amount of data. The integration of these data is important to understand the biological processes involved during bacterial colonization of the newborns intestine, particularly through breast milk. This work aims to exploit the bioinformatics approaches, to provide a new representation and interpretation of the interactions between differentially expressed genes in the host intestine induced by the microbiota.
- Research Article
74
- 10.1016/s0167-7799(99)01342-6
- Sep 1, 1999
- Trends in Biotechnology
Kleisli: a new tool for data integration in biology
- Research Article
3
- 10.2174/1574893615999210101125442
- Aug 11, 2021
- Current Bioinformatics
Integrating heterogeneous biological databases for unveiling the new intra-molecular and inter-molecular attributes, behaviors, and relationships in the human cellular system has always been a focused research area of computational biology. In this context, a lot of biological data integration systems have been deployed in the last couple of decades. One of the prime and common objectives of all these systems is to better facilitate the end-users for exploring, exploiting, and analyzing the integrated biological data for knowledge extraction. With the advent of especially high-throughput data generation technologies, biological data is growing and dispersing continuously, exponentially, heterogeneously, and geographically. Due to this, biological data integration systems face data integration and data organization-related current and future challenges. The objective of this review is to quantitatively evaluate and compare some of the recent warehouse- based multi-omics data integration systems to check their compliance with the current and future data integration needs. For this, we identified some of the major data integration design characteristics that should be in the multi-omics data integration model to comprehensively address the current and future data integration challenges. Based on these design characteristics and the evaluation criteria, we evaluated some of the recent data warehouse systems and showed categorical and comparative analysis results. Results show that most of the systems exhibit no or partial compliance with the required data integration design characteristics. So, these systems need design improvements to adequately address the current and future data integration challenges while keeping their service level commitments in place.
- Research Article
3
- 10.1089/cmb.2017.0199
- Apr 11, 2018
- Journal of Computational Biology
Life science studies represent one of the biggest generators of large data sets, mainly because of rapid sequencing technological advances. Biological networks including interactive networks and human curated pathways are essential to understand these high-throughput data sets. Biological network analysis offers a method to explore systematically not only the molecular complexity of a particular disease but also the molecular relationships among apparently distinct phenotypes. Currently, several packages for Python community have been developed, such as BioPython and Goatools. However, tools to perform comprehensive network analysis and visualization are still needed. Here, we have developed PyPathway, an extensible free and open source Python package for functional enrichment analysis, network modeling, and network visualization. The network process module supports various interaction network and pathway databases such as Reactome, WikiPathway, STRING, and BioGRID. The network analysis module implements overrepresentation analysis, gene set enrichment analysis, network-based enrichment, and de novo network modeling. Finally, the visualization and data publishing modules enable users to share their analysis by using an easy web application. For package availability, see the first Reference.
- Research Article
2
- 10.1002/cfg.134
- Feb 1, 2002
- Comparative and Functional Genomics
Beyond an analysis of gene function at the genome level, the ultimate goal of functional genomics is to understand the organisation and coordinated operation of the cell. Integration of data and information is an essential feature of all the steps leading from the production of experimental results to the modelling of a complete cell. The Geneva workshop, organised within the framework of the European Science Foundation (ESF) Programme on Integrated Approaches for Functional Genomics (http://www.functionalgenomics.org.uk), provided an excellent opportunity to review the issue of data integration from various perspectives. It brought together scientists with different backgrounds (biologists, bioinformaticians), who currently participate in consortia involving or requiring the integration of heterogeneous biological data. The workshop dedicated equal time to presentations and discussions in order to optimise knowledge distribution and sharing. The first session was devoted to existing functional genomics projects, where diverse sets of experimental approaches are applied to a model organism or given project. For these, both the scientific objectives and data management strategies were described. Another set of presentations focused on ‘data integration requirements’ in proteomics. In this rapidly evolving area, data integration is a key issue for technological developments that are being carried out in both academic and industrial contexts. Data integration approaches were then presented by scientists involved in the design and maintenance of public database resources. The last session was devoted to bioinformatics and gave an overview of state of the art approaches for building up information systems with good capabilities with respect to data integration. The workshop concluded with an open discussion on bottlenecks and perspectives that emerged from the presentations. As a number of speakers have submitted reviews based on their presentations that are appearing in this issue of CFG and in the next one, in this report Comparative and Functional Genomics Comp Funct Genom 2002; 3: 16–21. DOI: 10.1002 / cfg.134
- Book Chapter
4
- 10.4018/978-1-60566-308-1.ch012
- Jan 1, 2009
There is a proliferation of research and industrial organizations that produce sources of huge amounts of biological data issuing from experimentation with biological systems. In order to make these heterogeneous data sources easy to use, several efforts at data integration are currently being undertaken based mainly on XML. Starting from a discussion of the main biological data types and system interactions that need to be represented, the authors deal with the main approaches proposed for their modelling through XML. Then, they show the current efforts in biological data integration and how an increasing amount of Semantic information is required in terms of vocabulary control and ontologies. Finally, future research directions in biological data integration are discussed.
- Research Article
19
- 10.1145/2567663
- May 1, 2014
- Journal of Data and Information Quality
Data integration aims to combine heterogeneous information sources and to provide interfaces for accessing the integrated resource. Data integration is a collaborative task that may involve many people with different degrees of experience, knowledge of the application domain, and expectations relating to the integrated resource. It may be difficult to determine and control the quality of an integrated resource due to these factors. In this article, we propose a data integration methodology that has embedded within it iterative quality assessment and improvement of the integrated resource. We also propose an architecture for the realisation of this methodology. The quality assessment is based on an ontology representation of different users’ quality requirements and of the main elements of the integrated resource. We use description logic as the formal basis for reasoning about users’ quality requirements and for validating that an integrated resource satisfies these requirements. We define quality factors and associated metrics which enable the quality of alternative global schemas for an integrated resource to be assessed quantitively, and hence the improvement which results from the refinement of a global schema following our methodology to be measured. We evaluate our approach through a large-scale real-life case study in biological data integration in which an integrated resource is constructed from three autononous proteomics data sources.
- Research Article
187
- 10.1093/bioinformatics/btx656
- Oct 23, 2017
- Bioinformatics
Integrative omics is a central component of most systems biology studies. Computational methods are required for extracting meaningful relationships across different omics layers. Various tools have been developed to facilitate integration of paired heterogenous omics data; however most existing tools allow integration of only two omics datasets. Furthermore, existing data integration tools do not incorporate additional steps of identifying sub-networks or communities of highly connected entities and evaluating the topology of the integrative network under different conditions. Here we present xMWAS, a software for data integration, network visualization, clustering, and differential network analysis of data from biochemical and phenotypic assays, and two or more omics platforms. https://kuppal.shinyapps.io/xmwas (Online) and https://github.com/kuppal2/xMWAS/ (R). kuppal2@emory.edu. Supplementary data are available at Bioinformatics online.
- Book Chapter
266
- 10.1007/978-1-60761-175-2_12
- Jan 1, 2009
Cytoscape is a general network visualization, data integration, and analysis software package. Its development and use has been focused on the modeling requirements of systems biology, though it has been used in other fields. Cytoscape's flexibility has encouraged many users to adopt it and adapt it to their own research by using the plugin framework offered to specialize data analysis, data integration, or visualization. Plugins represent collections of community-contributed functionality and can be used to dynamically extend Cytoscape functionality. This community of users and developers has worked together since Cytoscape's initial release to improve the basic project through contributions to the core code and public offerings of plugin modules. This chapter discusses what Cytoscape does, why it was developed, and the extensions numerous groups have made available to the public. It also describes the development of a plugin used to investigate a particular research question in systems biology and walks through an example analysis using Cytoscape.