Accelerate Literature Icon
Want to do a literature review? Try our new Literature Review workflow

The harmonizome: a collection of processed datasets gathered to serve and mine knowledge about genes and proteins.

  • Abstract
  • Highlights & Summary
  • PDF
  • Literature Map
  • Similar Papers
Abstract
Translate article icon Translate Article Star icon

Genomics, epigenomics, transcriptomics, proteomics and metabolomics efforts rapidly generate a plethora of data on the activity and levels of biomolecules within mammalian cells. At the same time, curation projects that organize knowledge from the biomedical literature into online databases are expanding. Hence, there is a wealth of information about genes, proteins and their associations, with an urgent need for data integration to achieve better knowledge extraction and data reuse. For this purpose, we developed the Harmonizome: a collection of processed datasets gathered to serve and mine knowledge about genes and proteins from over 70 major online resources. We extracted, abstracted and organized data into ∼72 million functional associations between genes/proteins and their attributes. Such attributes could be physical relationships with other biomolecules, expression in cell lines and tissues, genetic associations with knockout mouse or human phenotypes, or changes in expression after drug treatment. We stored these associations in a relational database along with rich metadata for the genes/proteins, their attributes and the original resources. The freely available Harmonizome web portal provides a graphical user interface, a web service and a mobile app for querying, browsing and downloading all of the collected data. To demonstrate the utility of the Harmonizome, we computed and visualized gene–gene and attribute–attribute similarity networks, and through unsupervised clustering, identified many unexpected relationships by combining pairs of datasets such as the association between kinase perturbations and disease signatures. We also applied supervised machine learning methods to predict novel substrates for kinases, endogenous ligands for G-protein coupled receptors, mouse phenotypes for knockout genes, and classified unannotated transmembrane proteins for likelihood of being ion channels. The Harmonizome is a comprehensive resource of knowledge about genes and proteins, and as such, it enables researchers to discover novel relationships between biological entities, as well as form novel data-driven hypotheses for experimental validation.Database URL: http://amp.pharm.mssm.edu/Harmonizome.

Similar Papers
  • Preprint Article
  • 10.7490/f1000research.1113435.1
The harmonizome: a collection of processed datasets gathered to serve and mine knowledge about genes and proteins
  • Nov 21, 2016
  • F1000Research
  • Avi Ma'Ayan

Genomics, epigenomics, transcriptomics, proteomics and metabolomics efforts rapidly generate a plethora of data on the activity and levels of biomolecules within mammalian cells. At the same time, curation projects that organize knowledge from the biomedical literature into online databases are expanding. Hence, there is a wealth of information about genes, proteins and their associations, with an urgent need for data integration to achieve better knowledge extraction and data reuse. For this purpose, we developed the Harmonizome: a collection of processed datasets gathered to serve and mine knowledge about genes and proteins from over 70 major online resources. We extracted, abstracted and organized data into ∼72 million functional associations between genes/proteins and their attributes. Such attributes could be physical relationships with other biomolecules, expression in cell lines and tissues, genetic associations with knockout mouse or human phenotypes, or changes in expression after drug treatment. We stored these associations in a relational database along with rich metadata for the genes/proteins, their attributes and the original resources. The freely available Harmonizome web portal provides a graphical user interface, a web service and a mobile app for querying, browsing and downloading all of the collected data. To demonstrate the utility of the Harmonizome, we computed and visualized gene-gene and attribute-attribute similarity networks, and through unsupervised clustering, identified many unexpected relationships by combining pairs of datasets such as the association between kinase perturbations and disease signatures. We also applied supervised machine learning methods to predict novel substrates for kinases, endogenous ligands for G-protein coupled receptors, mouse phenotypes for knockout genes, and classified unannotated transmembrane proteins for likelihood of being ion channels. The Harmonizome is a comprehensive resource of knowledge about genes and proteins, and as such, it enables researchers to discover novel relationships between biological entities, as well as form novel data-driven hypotheses for experimental validation. URL: http://amp.pharm.mssm.edu/Harmonizome.

  • Peer Review Report
  • 10.7554/elife.70763.sa0
Editor's evaluation: Comparative transcriptomic analysis reveals translationally relevant processes in mouse models of malaria
  • Aug 11, 2021
  • Urszula Krzych

Comparative transcriptomics of whole blood can be used to evaluate the systemic host response and its concordance between human and mouse malaria and aid the selection of appropriate models for translational malaria research.

  • PDF Download Icon
  • Research Article
  • Cite Count Icon 2
  • 10.1007/s00213-024-06668-9
Late development of OCD-like phenotypes in Dlgap1 knockout mice
  • Aug 23, 2024
  • Psychopharmacology
  • Kimino Minagawa + 5 more

RationaleDespite variants in the Dlgap1 gene having the two lowest p-value in a genome-wide association study of obsessive compulsive disorder (OCD), previous studies reported the absence of OCD-like phenotypes in Dlgap1 knockout (KO) mice. Since these studies observed behavioral phenotypes only for a short period, development of OCD-like phenotypes in these mice at older ages was still plausible.ObjectiveTo examine the presence or absence of development of OCD-like phenotypes in Dlgap1 KO mice and their responsiveness to fluvoxamine.Methods and resultsNewly produced Dlgap1 KO mice were observed for a year. Modified SHIRPA primary screen in 2-month-old homozygous mutant mice showed only weak signs of anxiety, stress conditions and aggression. At older ages, however, these mutant mice exhibited excessive self-grooming characterized by increased scratching which led to skin lesions. A significant sex difference was observed in this scratching behavior. The penetrance of skin lesions reached 50% at 6–7 months of age and 90% at 12 months of age. In the open-field test performed just after the appearance of these lesions, homozygous mutant mice spent significantly less time in the center, an anxiety-like behavior, than did their wild-type and heterozygous littermates, none and less than 10% of which showed skin lesions at 1 year, respectively. The skin lesions and excessive self-grooming were significantly alleviated by two-week treatment with fluvoxamine.ConclusionUsefulness of Dlgap1 KO mice as a tool for investigating the pathogenesis of OCD-like phenotypes and its translational relevance was suggested.

  • Supplementary Content
  • 10.26267/unipi_dione/343
Πρόγραμμα πρώτων βοηθειών με χρήση της online βάσης δεδομένων Firebase και του web service Spring Boot
  • Nov 18, 2020
  • Dione (University of Piraeus)
  • Γεράσιμος Φλέγκας

The purpose of the present thesis is the development of a first aid system, that includes an Android application for the end user and a Web Service for the communication with the database. The application contains information on providing first aid regarding the emergency, as well as the capability of persisting the user’s location on the database in order for special assistance to arrive. The application does not communicate directly with the database, but the Web Service ensures this communication as an intermediate. For the development of the aforementioned software, full-stack technologies needed to be studied, meaning, ones that provide the capabilities of developing all functions, starting from the front-end, the user’s graphical interface, all the way to the back-end, the communication with the database and the data transportation. For the user part, as mentioned, an Android application was decided to be developed. The Web Service software uses the REST architecture for the communicaton with the database and lastly for the storage of the data the Google Firebase platform was chosen, which among others services, provides a non relational database. The data are saved in JSON format and are available in real time.

  • Research Article
  • Cite Count Icon 20
  • 10.2174/1568007024606203
Novel directions in antipsychotic target identification using gene arrays.
  • Apr 1, 2002
  • Current drug targets. CNS and neurological disorders
  • Michael Palfreyman + 5 more

Schizophrenia is a major health problem that affects 2 million individuals in the United States. Antipsychotics offer considerable symptomatic relief and, although commonly discovered by screening with single biological targets, most interact with multiple receptors and signaling pathways. Considerable evidence from family and twin studies demonstrates genetic components and multiple chromosomal regions associated with schizophrenia. The polygenic nature of schizophrenia and multiple mechanisms for most effective agents indicate the need for broader approaches to target identification. Gene expression profiling of post-mortem human brain tissue simultaneously reveals the expression of many thousands of genes. A comparison of tissue from normals and patients provides a 'disease signature' of aberrantly expressed genes. 'Drug signatures' are the gene expression changes of cultured human or animal neurons treated with psychiatric drugs, and from animals chronically treated with these drugs. A selection of genes from disease and drug signatures can create a set of targets whose changes may better predict disease and its treatment by effective agents. This multi-parameter high throughput screening (MPHTS(SM)) approach evaluates the mRNA expression pattern of cultured cells exposed to candidate compounds. Compounds that normalize genes altered in schizophrenia may better address its underlying causes. Drugs that mimic gene expression changes that are consistently altered by effective antipsychotic agents provide a drug improvement strategy if efficacy is enhanced or side effects are attenuated.

  • Conference Article
  • Cite Count Icon 19
  • 10.1109/services-2.2008.23
An Adaptive User Interface Generation Framework for Web Services
  • Sep 1, 2008
  • Jiang He + 4 more

In a service chain, one web service invokes another based on the WSDL definition. For some Web services, the invoker may be a user. The ways for a service to interact with an application and a user should be different. When the service interacts with a user, it is preferable to provide a friendly graphical user interface (GUI). Also, users may access a web service from different devices, such as desktops and various handheld devices, and the GUI should be designed differently to fit different characteristics of the devices. However, it would be a burden to the web service designer to manually develop the interfaces for potential user interactions or even manually design various interfaces to fit potential user devices. In this paper, an adaptive user interface generation framework is presented to automate the task. The framework takes the WSDL interface specification as input to generate customized GUI for the specific user device. Additional specifications, such as the OWL specification for interface object semantics, can be provided to guide the generation of better GUI.The framework generates GUI in two stages. First, based on input specifications and device characteristics, it generates display preferences and constrains. Then, it invokes a decision algorithm to make the optimal GUI layout decision that satisfies all the display constrains and maximizes the preference values. A case study has been explored and it shows that the framework can effectively and adaptively generate appropriate GUIs for various user devices.

  • Conference Article
  • Cite Count Icon 7
  • 10.1109/scc.2006.118
Web Services on Rails: Using Ruby and Rails for Web Services Development and Mashups
  • Sep 1, 2006
  • E Michael Maximilien

One of the interesting aspects of the Web 2.0 'evolution' is the wide-availability of various Web applications as APIs or Web services. These APIs expose informational services on the Web and take many forms of remote invocation of functions using standard Web protocols and XML for data representations, e.g., REST, SOAP/WSDL, XML-RPC, and other approaches. The services (or APIs) are also usually accompanied by user facing Web applications for human-consumption. Canonical examples are Google Maps, Yahoo! Flykr and del.icio.us, EVDB's Eventful's application and API, Amazon.com's S3, ECS, Alexa, and many others. The Ruby programming language and its Rails framework are ideal for programming Web applications and services in the Web 2.0. Ruby's modern and dynamic features make it an excellent language for rapid prototyping and integration of various Web services. Rails' superb support for rapid Web application development, database access, and AJAX, make it well suited for creating front-ends and back-ends to the next generation of Web applications and services. In this tutorial we will take a hands-on deep-dive into the Ruby and Rails platform and learn how they can be used to: (1) create Web applications backed by a relational database, (2) consume Web services, (3) create and deploy APIs or Web services, and (4) mashup of existing Web services and applications. No a priori knowledge of Ruby or Rails is required - although some programming in a modern OO language and Web application development are definite plus.

  • Research Article
  • Cite Count Icon 247
  • 10.1194/jlr.m500362-jlr200
One generation of n-3 polyunsaturated fatty acid deprivation increases depression and aggression test scores in rats
  • Jan 1, 2006
  • Journal of Lipid Research
  • James C Demar + 5 more

Male rat pups at weaning (21 days of age) were subjected to a diet deficient or adequate in n-3 polyunsaturated fatty acids (n-3 PUFAs) for 15 weeks. Performance on tests of locomotor activity, depression, and aggression was measured in that order during the ensuing 3 weeks, after which brain lipid composition was determined. In the n-3 PUFA-deprived rats, compared with n-3 PUFA-adequate rats, docosahexaenoic acid (22:6n-3) in brain phospholipid was reduced by 36% and docosapentaenoic acid (22:5n-6) was elevated by 90%, whereas brain phospholipid concentrations were unchanged. N-3 PUFA-deprived rats had a significantly increased (P = 0.03) score on the Porsolt forced-swim test for depression, and increased blocking time (P = 0.03) and blocking number (P = 0.04) scores (uncorrected for multiple comparisons) on the isolation-induced resident-intruder test for aggression. Large effect sizes (d > 0.8) were found on the depression score and on the blocking time score of the aggression test. Scores on the open-field test for locomotor activity did not differ significantly between groups, and had only small to medium effect sizes. This single-generational n-3 PUFA-deprived rat model, which demonstrated significant changes in brain lipid composition and in test scores for depression and aggression, may be useful for elucidating the contribution of disturbed brain PUFA metabolism to human depression, aggression, and bipolar disorder.

  • Conference Article
  • Cite Count Icon 2
  • 10.1109/carpathiancc.2013.6560574
Graphic presentation of selected thermochemical properties of substances in Matlab using service oriented architecture
  • May 1, 2013
  • Jan Terpak + 3 more

This contribution deals with the disclosure of the thermochemical properties of individual substances by using a service-oriented architecture through a web service that is implemented in the client environment MATLAB. The basic principles of service-oriented architecture and its use by web service are described and a specific web service that provides different thermochemical properties such as molar heat capacity, enthalpy, entropy, and the like is listed. The following describes the design and implementation of functions for the MATLAB client environment, which makes it easy to obtain, to elaborate and to make available the thermochemical properties provided by the web service. Solution of application interface is based on taking advantage of graphic user interface support MATLAB and mentioned web service. Application interface consists of parameters for identification of substance, temperature's scope and step, graph size and x-axis ticketing density as well as if appropriate of file name in which calculated graph course is in graphics format. Finally, specific examples are shown using created functions and graphics interface. Possibilities of potential uses are discussed.

  • Preprint Article
  • 10.7490/f1000research.1119701.1
Interactive client tools and galaxy ProTo 2.0 redux - Empowering researchers to implement a GUI alongside methodology by using Galaxy as an IDE
  • Jan 31, 2024
  • Faculty of 1000 Research Ltd
  • Sveinung Gundersen + 1 more

The Galaxy platform is one of a select few successful attempts at creating graphical user interfaces (GUI) for analyses within bioinformatics. Still, the additional cost for researchers to create GUIs for their methodological contributions or software tools is often found to be too high, as also witnessed in a recent Twitter discussion that received some attention. One unspoken premise for the discussion on the cost/benefit of developing GUIs is this: one will always first develop a tool for command line or programmatic use, and then, if popular, one might start to consider implementing a GUI. Contrary to this premise, we have years of experience within a research environment that tells us there is great benefit (and also some pitfalls) to be found by switching this order and instead start directly on developing tools and methodologies with GUIs. Moreover, we have found Galaxy to be an excellent platform for prototypical, GUI-first development of tools. In the Galaxy Community Conferences (GCC) in 2015 and 2016 we presented Galaxy ProTo v1 which demonstrated a novel approach to tool development in Galaxy. By replacing the XML middle layer with a template-based API in Python, we provided researchers, master students and other tool developers with a platform for on-the-fly development of highly dynamic web tools using pure Python. The GUI contents were defined by passing appropriately structured data to the GUI-generating templates in Galaxy ProTo. Based on data structures like strings, lists, dicts, and lists of lists, the GUI templates would automatically and immediately generate corresponding user interface elements: text boxes, drop-down lists, multiple selection list boxes, and tables, respectively. By keeping GUI development simple and intuitive and without the overhead of mastering the details of traditional Galaxy tool development, ProTo v1 tools were ideal for prototyping novel functionality as demonstrators to supervisors or providing early access for research collaborators to test out "avant garde" functionality,. To this end, vanilla Galaxy also provides powerful sharing features, the automatic provenance of Galaxy histories, as well as the rerun and bug reporting functionalities, which can all be powerfully used as part of a development cycle focused on immediate delivery to end users (which might include yourself). Unfortunately, Galaxy ProTo v1 was only available as an unofficial fork of the Galaxy code base and the functionality was never offered as part of a vanilla Galaxy installation. After lying dormant for a few years, we have now resurrected the idea and code base but replaced the technological underpinnings to allow for distribution and use with vanilla Galaxy installations. Galaxy ProTo v2 runs as a web server in a Docker (or Podman) container as a separate, long-running Galaxy Interactive Tools job that also connects to the Galaxy API. Through a newly proposed tool type, Interactive Client Tools, other tools can connect to this dockerized web server, which then generates the dynamic GUI content of Galaxy ProTo , displayed within the middle panel of Galaxy. From the GUI, ProTo jobs can be submitted for execution in Galaxy in the same way from standard tools. Furthermore, with the use of file synchronization software as a continuous deployment solution, we can once again provide on-the-fly development of highly dynamic prototyping tools, now in vanilla Galaxy. We will present an early technological demo of Galaxy ProTo v2 in order to receive community feedback on the functionality and raise interest from core developers in the Galaxy Projects to help integrate a few generally useful additions to the Galaxy code base that is needed to make Galaxy ProTo v2 work on a vanilla Galaxy installation.

  • PDF Download Icon
  • Research Article
  • Cite Count Icon 5
  • 10.1186/1471-2105-14-126
Clever generation of rich SPARQL queries from annotated relational schema: application to Semantic Web Service creation for biological databases
  • Apr 15, 2013
  • BMC Bioinformatics
  • Julien Wollbrett + 3 more

BackgroundIn recent years, a large amount of “-omics” data have been produced. However, these data are stored in many different species-specific databases that are managed by different institutes and laboratories. Biologists often need to find and assemble data from disparate sources to perform certain analyses. Searching for these data and assembling them is a time-consuming task. The Semantic Web helps to facilitate interoperability across databases. A common approach involves the development of wrapper systems that map a relational database schema onto existing domain ontologies. However, few attempts have been made to automate the creation of such wrappers.ResultsWe developed a framework, named BioSemantic, for the creation of Semantic Web Services that are applicable to relational biological databases. This framework makes use of both Semantic Web and Web Services technologies and can be divided into two main parts: (i) the generation and semi-automatic annotation of an RDF view; and (ii) the automatic generation of SPARQL queries and their integration into Semantic Web Services backbones. We have used our framework to integrate genomic data from different plant databases.ConclusionsBioSemantic is a framework that was designed to speed integration of relational databases. We present how it can be used to speed the development of Semantic Web Services for existing relational biological databases. Currently, it creates and annotates RDF views that enable the automatic generation of SPARQL queries. Web Services are also created and deployed automatically, and the semantic annotations of our Web Services are added automatically using SAWSDL attributes. BioSemantic is downloadable at http://southgreen.cirad.fr/?q=content/Biosemantic.

  • Research Article
  • Cite Count Icon 1
  • 10.1016/j.placenta.2023.05.012
Novel genes associated with a placental phenotype in knockout mice also respond to cellular stressors in primary human trophoblasts
  • May 19, 2023
  • Placenta
  • Elif Kadife + 4 more

Novel genes associated with a placental phenotype in knockout mice also respond to cellular stressors in primary human trophoblasts

  • PDF Download Icon
  • Research Article
  • Cite Count Icon 45
  • 10.1371/journal.pone.0014319
Identification of Genes and Networks Driving Cardiovascular and Metabolic Phenotypes in a Mouse F2 Intercross
  • Dec 14, 2010
  • PLoS ONE
  • Jonathan M J Derry + 17 more

To identify the genes and pathways that underlie cardiovascular and metabolic phenotypes we performed an integrated analysis of a mouse C57BL/6J x A/J F2 (B6AF2) cross by relating genome-wide gene expression data from adipose, kidney, and liver tissues to physiological endpoints measured in the population. We have identified a large number of trait QTLs including loci driving variation in cardiac function on chromosomes 2 and 6 and a hotspot for adiposity, energy metabolism, and glucose traits on chromosome 8. Integration of adipose gene expression data identified a core set of genes that drive the chromosome 8 adiposity QTL. This chromosome 8 trans eQTL signature contains genes associated with mitochondrial function and oxidative phosphorylation and maps to a subnetwork with conserved function in humans that was previously implicated in human obesity. In addition, human eSNPs corresponding to orthologous genes from the signature show enrichment for association to type II diabetes in the DIAGRAM cohort, supporting the idea that the chromosome 8 locus perturbs a molecular network that in humans senses variations in DNA and in turn affects metabolic disease risk. We functionally validate predictions from this approach by demonstrating metabolic phenotypes in knockout mice for three genes from the trans eQTL signature, Akr1b8, Emr1, and Rgs2. In addition we show that the transcriptional signatures for knockout of two of these genes, Akr1b8 and Rgs2, map to the F2 network modules associated with the chromosome 8 trans eQTL signature and that these modules are in turn very significantly correlated with adiposity in the F2 population. Overall this study demonstrates how integrating gene expression data with QTL analysis in a network-based framework can aid in the elucidation of the molecular drivers of disease that can be translated from mice to humans.

  • PDF Download Icon
  • Research Article
  • Cite Count Icon 31
  • 10.1074/jbc.m413806200
Platelet-derived Growth Factor Stimulates Src-dependent mRNA Stabilization of Specific Early Genes in Fibroblasts
  • Mar 1, 2005
  • Journal of Biological Chemistry
  • Paul A Bromann + 5 more

The Src family of protein-tyrosine kinases (SFKs) participates in a variety of signal transduction pathways, including promotion of cell growth, prevention of apoptosis, and regulation of cell interactions and motility. In particular, SFKs are required for the mitogenic response to platelet-derived growth factor (PDGF). However, it is not clear whether there is a discrete SFK-specific pathway leading to enhanced gene expression or whether SFKs act to generally enhance PDGF-stimulated gene expression. To examine this, we treated quiescent NIH3T3 cells with PDGF in the presence or absence of small molecule inhibitors of SFKs, phosphatidylinositol 3-kinase (PI3K), and MEK1/2. Global patterns of gene expression were analyzed by using Affymetrix Gene-Chip arrays, and data were validated by using reverse transcription-PCR and ribonuclease protection assay. We identified a discrete set of immediate early genes induced by PDGF and inhibited in the presence of the SFK-selective inhibitor SU6656. A subset of these SFK-dependent genes was induced by PDGF even in the presence of the MEK1/2 inhibitor U0126 or the PI3K inhibitor LY294002. By using ribonuclease protection assays and nuclear run-off assays, we further determined that PDGF did not stimulate the rate of transcription of these SFK-dependent immediate early genes but rather promoted mRNA stabilization. Our data suggest that PDGF regulates gene expression through an SFK-specific pathway that is distinct from the Ras-MAPK and PI3K pathways, and that SFKs signal gene expression by enhancing mRNA stability.

  • Conference Article
  • Cite Count Icon 5
  • 10.1109/ds-rt.2011.20
Maturing Supporting Software for C2-Simulation Interoperation
  • Sep 1, 2011
  • J Mark Pullen + 1 more

The Battle Management Language (BML) has been developed as an unambiguous representation of orders, reports and requests between military command and control (C2) systems and simulations. This paper describes development of two critical elements for BML experimentation. The first is a Web service that is used as a repository for the orders, reports, and requests, packaged as XML documents and stored in a relational database in a standard format. The second is a graphic user interface that supports inspection and modification of BML documents using forms that are created from the BML schemas at runtime, combined with a geospatial interface. This paper describes the role of these elements in C2-simulation interoperation and describes their architecture and features, as motivated by application in NATO Coalition BML experimentation. The software described is available as open source.

Save Icon
Up Arrow
Open/Close
Setting-up Chat
Loading Interface