Data Flow Operations Research Articles

Parallel dataflow engines such as Apache Hadoop, Apache Spark, and Apache Flink are an established alternative to relational databases for modern data analysis applications. A characteristic of these systems is a scalable programming model based on distributed collections and parallel transformations expressed by means of second-order functions such as map and reduce. Notable examples are Flink’s DataSet and Spark’s RDD programming abstractions. These programming models are realized as EDSLs—domain specific languages embedded in a general-purpose host language such as Java, Scala, or Python. This approach has several advantages over traditional external DSLs such as SQL or XQuery. First, syntactic constructs from the host language (e.g., anonymous functions syntax, value definitions, and fluent syntax via method chaining) can be reused in the EDSL. This eases the learning curve for developers already familiar with the host language. Second, it allows for seamless integration of library methods written in the host language via the function parameters passed to the parallel dataflow operators. This reduces the effort for developing analytics dataflows that go beyond pure SQL and require domain-specific logic. At the same time, however, state-of-the-art parallel dataflow EDSLs exhibit a number of shortcomings. First, one of the main advantages of an external DSL such as SQL—the high-level, declarative Select-From-Where syntax—is either lost completely or mimicked in a non-standard way. Second, execution aspects such as caching, join order, and partial aggregation have to be decided by the programmer. Optimizing them automatically is very difficult due to the limited program context available in the intermediate representation of the DSL. In this article, we argue that the limitations listed above are a side effect of the adopted type-based embedding approach. As a solution, we propose an alternative EDSL design based on quotations. We present a DSL embedded in Scala and discuss its compiler pipeline, intermediate representation, and some of the enabled optimizations. We promote the algebraic type of bags in union representation as a model for distributed collections and its associated structural recursion scheme and monad as a model for parallel collection processing. At the source code level, Scala’s comprehension syntax over a bag monad can be used to encode Select-From-Where expressions in a standard way. At the intermediate representation level, maintaining comprehensions as a first-class citizen can be used to simplify the design and implementation of holistic dataflow optimizations that accommodate for nesting and control-flow. The proposed DSL design therefore reconciles the benefits of embedded parallel dataflow DSLs with the declarativity and optimization potential of external DSLs like SQL.

Read full abstract

The tremendous increase in the computing capacity of the embedded architectures has led to widespread deployment of embedded applications. These applications generally exhibit similar patterns in their specification such as filters in which multiply and accumulate operations are repetitive. If such patterns are identified and used for the system design, trade-off between the area and delay can be achieved. This paper proposes a new methodology which allows to implements a design by retrieving similar patterns known as graph isomorphs and interfaces them as HW accelerators in the system-on-chip design flow. An effective algorithm that converges in polynomial time has been proposed to find such similar subgraphs. In the next phase of the design flow, an algorithm has been proposed which performs the scheduling of clusters and minimizes the time overhead. All algorithms have been written in python for parsing the data flow description and test the correctness of the proposed work. The proposed design flow has been applied to five different programs which are sine, cosine, exponent, matrix multiplication and discrete cosine transform (DCT). These have been described as a data flow graph and have been used for results comparison. An estimation table showing the HW and SW parameter of the data flow operators has been developed for timing and area analysis of the programs. The work is an effort to show the clustering and scheduling of a standalone specification which is mapped on static reconfigurable fabric. Reconfigurable computing systems (RCS) are a popular platform for embedded computing applications as they offer a wide exploration in the design space by allowing HW, SW or HW–SW (hybrid) implementation depending on computational demand and resource requirement. These systems have inspired the designers to find new frameworks for achieving the optimized system characteristics under the given constraints. Any static or dynamic HW hardware optimization in an application can be proposed, implemented and easily verified on the chip. The results presented show the comparison of the proposed approach with SW and HW implementation of DCT design on the Xilinx ML507 board. HW timer has been used to find the execution of each implementation. The experimental verification of the proposed algorithms shows that static IP core design flow gives better results.

Read full abstract

Data Flow Operations Research Articles

Articles published on Data Flow Operations

FlowCert: Translation Validation for Asynchronous Dataflow via Dynamic Fractional Permissions

Query Planning for Robust and Scalable Hybrid Network Telemetry Systems

A Streaming Data Processing Architecture Based on Lookup Tables

Hybrid Cloud Architecture for Higher Education System

Scalable Linear Algebra on a Relational Database System

Foreword to the special issue of the workshop on high performance computing systems (XVIII Simpósio em Sistemas Computacionais de Alto Desempenho, WSCAD 2017)

Representations and Optimizations for Embedded Parallel Dataflow Languages

ADD: Accelerator Design and Deploy ‐ A tool for FPGA high‐performance dataflow computing

Using graph isomorphism for mapping of data flow applications on reconfigurable computing systems

Explaining outputs in modern data analytics

Scalable and adaptive online joins

Nesting Strategies for Enabling Nimble MapReduce Dataflows for Large RDF Data

A data‐operation model based on partial vector space for batch processing in workflow

Double Input Operators of the DF KPI System

A Parallel Prolog Abstract Machine and its Multi-Transputer Implementation

Automated Design of Circuits from Recursion Equations Using Theorem‐Proving Technique

Process and dataflow control in distributed data-intensive systems

The Effect of Operation Scheduling on the Performance of a Data Flow Computer

New VLSI systolic array design for real-time digital signal processing

EPILOG = PROLOG + Data Flow

Lead the way for us

Editage

Paperpal

R Discovery

Mind the Graph

Data Flow Operations Research Articles

Articles published on Data Flow Operations

FlowCert: Translation Validation for Asynchronous Dataflow via Dynamic Fractional Permissions

Query Planning for Robust and Scalable Hybrid Network Telemetry Systems

A Streaming Data Processing Architecture Based on Lookup Tables

Hybrid Cloud Architecture for Higher Education System

Scalable Linear Algebra on a Relational Database System

Foreword to the special issue of the workshop on high performance computing systems (XVIII Simpósio em Sistemas Computacionais de Alto Desempenho, WSCAD 2017)

Representations and Optimizations for Embedded Parallel Dataflow Languages

ADD: Accelerator Design and Deploy ‐ A tool for FPGA high‐performance dataflow computing

Using graph isomorphism for mapping of data flow applications on reconfigurable computing systems

Explaining outputs in modern data analytics

Scalable and adaptive online joins

Nesting Strategies for Enabling Nimble MapReduce Dataflows for Large RDF Data

A data‐operation model based on partial vector space for batch processing in workflow

Double Input Operators of the DF KPI System

A Parallel Prolog Abstract Machine and its Multi-Transputer Implementation

Automated Design of Circuits from Recursion Equations Using Theorem‐Proving Technique

Process and dataflow control in distributed data-intensive systems

The Effect of Operation Scheduling on the Performance of a Data Flow Computer

New VLSI systolic array design for real-time digital signal processing

EPILOG = PROLOG + Data Flow