BioWarehouse: a bioinformatics database warehouse toolkit.

Thomas J Lee,Priyanka Gupta,Peter D Karp,Jessica D Tenenbaum,Yannick Pouliot,David Wj Stringer-Calvert,Valerie Wagner

doi:10.1186/1471-2105-7-170

Thomas J Lee, Priyanka Gupta + Show 5 more

Open Access

https://doi.org/10.1186/1471-2105-7-170

Copy DOI

Abstract

BackgroundThis article addresses the problem of interoperation of heterogeneous bioinformatics databases.ResultsWe introduce BioWarehouse, an open source toolkit for constructing bioinformatics database warehouses using the MySQL and Oracle relational database managers. BioWarehouse integrates its component databases into a common representational framework within a single database management system, thus enabling multi-database queries using the Structured Query Language (SQL) but also facilitating a variety of database integration tasks such as comparative analysis and data mining. BioWarehouse currently supports the integration of a pathway-centric set of databases including ENZYME, KEGG, and BioCyc, and in addition the UniProt, GenBank, NCBI Taxonomy, and CMR databases, and the Gene Ontology. Loader tools, written in the C and JAVA languages, parse and load these databases into a relational database schema. The loaders also apply a degree of semantic normalization to their respective source data, decreasing semantic heterogeneity. The schema supports the following bioinformatics datatypes: chemical compounds, biochemical reactions, metabolic pathways, proteins, genes, nucleic acid sequences, features on protein and nucleic-acid sequences, organisms, organism taxonomies, and controlled vocabularies. As an application example, we applied BioWarehouse to determine the fraction of biochemically characterized enzyme activities for which no sequences exist in the public sequence databases. The answer is that no sequence exists for 36% of enzyme activities for which EC numbers have been assigned. These gaps in sequence data significantly limit the accuracy of genome annotation and metabolic pathway prediction, and are a barrier for metabolic engineering. Complex queries of this type provide examples of the value of the data warehousing approach to bioinformatics research.ConclusionBioWarehouse embodies significant progress on the database integration problem for bioinformatics.

Highlights

This article addresses the problem of interoperation of heterogeneous bioinformatics databases
We present results obtained by BioWarehouse in its use by several bioinformatics projects, and a performance analysis of BioWarehouse
An SRI project is developing algorithms for predicting which genes within a sequenced genome code for missing enzymes within metabolic pathways predicted for that genome [29]

Summary

Introduction

This article addresses the problem of interoperation of heterogeneous bioinformatics databases. One approach has involved mediator-based solutions that transmit multidatabase queries to multiple source DBs across the Internet. Some progress has been made in developing mediator technology, we argue that these systems face several practical limitations (see Section "Comparison of the Warehouse and Multidatabase Approaches" for more details), including that (a) few source DBs accept complex queries via the Internet (an immediate deal killer), (b) the user lacks control over which version of the data is queried, and over the hardware that provides query processing power, (c) the speed of the Internet limits transmission of query results, and (d) users cannot cleanse the source DBs that they query of potentially erroneous, incomplete or redundant data – that is, they cannot alter the source DBs in any way.

Objectives

Methods

Results

Conclusion

Talk to us

Join us for a 30 min session where you can share your feedback and ask us any queries you have

Schedule a call

Journal: BMC Bioinformatics	Publication Date: Mar 23, 2006
Citations: 196	License type: cc-by

R Discovery Prime

R Discovery Prime

BioWarehouse: a bioinformatics database warehouse toolkit.

Abstract

Highlights

Summary

Talk to us

Similar Papers

More From: BMC Bioinformatics

Lead the way for us

Similar Papers

The Edinburgh human metabolic network reconstruction and its functional analysis
Hongwu Ma ... Igor Goryanin
Molecular Systems Biology | VOL. 3
Hongwu Ma, et. al.Hongwu Ma ... Igor Goryanin
01 Jan 2007
Molecular Systems Biology | VOL. 3

A Mass Spectrometry Proteomics Data Management Platform
Vagisha Sharma ... Michael Riffle
Molecular & Cellular Proteomics | VOL. 11
Vagisha Sharma, et. al.Vagisha Sharma ... Michael Riffle
01 Sep 2012
Molecular & Cellular Proteomics | VOL. 11

Identifying Candidate Genes Using the BioWarehouse: A Case Study
Y Pouliot ... V Wagner
-
Y Pouliot, et. al.Y Pouliot ... V Wagner
16 Aug 2005
16 Aug 2005

On resolving schematic heterogeneity in multidatabase systems
Won Kim ... Mark Scheevel
Distributed and Parallel Databases | VOL. 1
Won Kim, et. al.Won Kim ... Mark Scheevel
01 Jul 1993
Distributed and Parallel Databases | VOL. 1

Editage

Paperpal

R Discovery

Mind the Graph

R Discovery Prime

R Discovery Prime

BioWarehouse: a bioinformatics database warehouse toolkit.

Abstract

Highlights

Summary

Talk to us

Similar Papers

More From: BMC Bioinformatics