Streaming data analytics via message passing with application to graph algorithms

Steven J Plimpton,Tim Shead

doi:10.1016/j.jpdc.2014.04.001

Abstract

The need to process streaming data, which arrives continuously at high-volume in real-time, arises in a variety of contexts including data produced by experiments, collections of environmental or network sensors, and running simulations. Streaming data can also be formulated as queries or transactions which operate on a large dynamic data store, e.g. a distributed database.We describe a lightweight, portable framework named PHISH which provides a communication model enabling a set of independent processes to compute on a stream of data in a distributed-memory parallel manner. Datums are routed between processes in patterns defined by the application. PHISH provides multiple communication backends including MPI and sockets/ZMQ. The former means streaming computations can be run on any parallel machine which supports MPI; the latter allows them to run on a heterogeneous, geographically dispersed network of machines.We illustrate how streaming MapReduce operations can be implemented using the PHISH communication model, and describe streaming versions of three algorithms for large, sparse graph analytics: triangle enumeration, sub-graph isomorphism matching, and connected component finding. We also provide benchmark timings comparing MPI and socket performance for several kernel operations useful in streaming algorithms.

Talk to us

Join us for a 30 min session where you can share your feedback and ask us any queries you have

Schedule a call

R Discovery Prime

R Discovery Prime

Streaming data analytics via message passing with application to graph algorithms

Abstract

Talk to us

Similar Papers

More From: Journal of Parallel and Distributed Computing

Lead the way for us

Journal: Journal of Parallel and Distributed Computing	Publication Date: May 6, 2014
Citations: 25

Similar Papers

Exploring Data Streams with Nonparametric Estimators
C Heinz ... B Seeger
-
C Heinz, et. al.C Heinz ... B Seeger
03 Jul 2006
03 Jul 2006

NStreamAware
Fabian Fischer ... Daniel A Keim
-
Fabian Fischer, et. al.Fabian Fischer ... Daniel A Keim
10 Nov 2014
10 Nov 2014

Geospatial Data Streams
Zdravko Galić
-
Zdravko GalićZdravko Galić
24 Oct 2017
24 Oct 2017

A real-time top-k query algorithm and parallelized implementation
Yao Lu ... Jun Liu
-
Yao Lu, et. al. Yao Lu ... Jun Liu
01 Nov 2014
01 Nov 2014

Editage

Paperpal

R Discovery

Mind the Graph

R Discovery Prime

R Discovery Prime

Streaming data analytics via message passing with application to graph algorithms

Abstract

Talk to us

Similar Papers

More From: Journal of Parallel and Distributed Computing