The BioDICE Taverna plugin for clustering and visualization of biological data: a workflow for molecular compounds exploration

Antonino Fiannaca,Riccardo Rizzo,Salvatore Gaglio,Giuseppe Di Fatta,Massimo La Rosa,Alfonso Urso

doi:10.1186/1758-2946-6-24

Antonino Fiannaca, Riccardo Rizzo + Show 4 more

Open Access

https://doi.org/10.1186/1758-2946-6-24

Copy DOI

Abstract

BackgroundIn many experimental pipelines, clustering of multidimensional biological datasets is used to detect hidden structures in unlabelled input data. Taverna is a popular workflow management system that is used to design and execute scientific workflows and aid in silico experimentation. The availability of fast unsupervised methods for clustering and visualization in the Taverna platform is important to support a data-driven scientific discovery in complex and explorative bioinformatics applications.ResultsThis work presents a Taverna plugin, the Biological Data Interactive Clustering Explorer (BioDICE), that performs clustering of high-dimensional biological data and provides a nonlinear, topology preserving projection for the visualization of the input data and their similarities. The core algorithm in the BioDICE plugin is Fast Learning Self Organizing Map (FLSOM), which is an improved variant of the Self Organizing Map (SOM) algorithm. The plugin generates an interactive 2D map that allows the visual exploration of multidimensional data and the identification of groups of similar objects. The effectiveness of the plugin is demonstrated on a case study related to chemical compounds.ConclusionsThe number and variety of available tools and its extensibility have made Taverna a popular choice for the development of scientific data workflows. This work presents a novel plugin, BioDICE, which adds a data-driven knowledge discovery component to Taverna. BioDICE provides an effective and powerful clustering tool, which can be adopted for the explorative analysis of biological datasets.

Highlights

In many experimental pipelines, clustering of multidimensional biological datasets is used to detect hidden structures in unlabelled input data
In previous works [6,8], the Fast Learning Self Organizing Map (FLSOM) algorithm was shown to be very effective in the cluster analysis of molecular compounds
The Biological Data Interactive Clustering Explorer (BioDICE) Taverna plugin requires an input file containing a features × patterns data table, that is a matrix with the feature identifiers as rows and the chemical

Summary

Results

This work presents a Taverna plugin, the Biological Data Interactive Clustering Explorer (BioDICE), that performs clustering of high-dimensional biological data and provides a nonlinear, topology preserving projection for the visualization of the input data and their similarities. The core algorithm in the BioDICE plugin is Fast Learning Self Organizing Map (FLSOM), which is an improved variant of the Self Organizing Map (SOM) algorithm. The plugin generates an interactive 2D map that allows the visual exploration of multidimensional data and the identification of groups of similar objects. The effectiveness of the plugin is demonstrated on a case study related to chemical compounds

Conclusions

Background

Results and discussion