Analysis of Blocking and Scheduling for FPGA-Based Floating-Point Matrix Multiplication Analyse du blocage et de l’ordonnancement d’une multiplication matricielle à virgule flottante sur un FPGA

Ahmad Khayyat,Naraig Manjikian

doi:10.1109/cjece.2014.2317983

Abstract

This paper considers blocking and scheduling for the design and implementation of field-programmable gate array (FPGA)-based floating-point parallel matrix multiplication in the presence of a memory hierarchy. For high performance, on-chip memory holds data that are reused when the computation is divided into blocks, and multiple arithmetic units perform independent operations within each block in parallel. The first contribution of this paper is a detailed analysis of the design space to characterize performance based on the amount of on-chip memory used and the approaches considered for blocking and scheduling of the computation. A comparison is also made to prior work with a unified view. The second contribution is a flexible high-performance implementation for the Altera Stratix IV EP4SGX530C2 FPGA with an interface to external double-data-rate synchronous dynamic RAM (DDR2 SDRAM) memory. Various configuration options support optimization of different objectives, and the resulting configurations have been verified in simulation and in hardware. For double-precision floating-point, a performance of 16 giga-floating-point operations per second (GFLOPS) is achievable with 64 arithmetic units at 160 MHz.

Talk to us

Join us for a 30 min session where you can share your feedback and ask us any queries you have

Schedule a call

R Discovery Prime

R Discovery Prime

Analysis of Blocking and Scheduling for FPGA-Based Floating-Point Matrix Multiplication Analyse du blocage et de l’ordonnancement d’une multiplication matricielle à virgule flottante sur un FPGA

Abstract

Talk to us

Similar Papers

More From: Canadian Journal of Electrical and Computer Engineering

Lead the way for us

Journal: Canadian Journal of Electrical and Computer Engineering	Publication Date: Jan 1, 2014
Citations: 4

Similar Papers

FPGA Implementation of an Evolving Spiking Neural Network
Alan Zuppicich ... Snjezana Soltic
-
Alan Zuppicich, et. al.Alan Zuppicich ... Snjezana Soltic
01 Jan 2009
01 Jan 2009

FPGA design for constrained energy minimization
Jianwei Wang ... Chein-I Chang
-
Jianwei Wang, et. al.Jianwei Wang ... Chein-I Chang
27 Feb 2004
27 Feb 2004

FPGA Design and Implementation of Kinect-Like Depth Sensing
Jiao Wang ... Feng Wu
IEEE Transactions on Circuits and Systems for Video Technology | VOL. 26
Jiao Wang, et. al.Jiao Wang ... Feng Wu
01 Jun 2016
IEEE Transactions on Circuits and Systems for Video Technology | VOL. 26

Python-based FPGA implementation of AES using Migen for Internet of Things Security
K O Setetemela ... K Keta
-
K O Setetemela, et. al.K O Setetemela ... K Keta
01 Feb 2019
01 Feb 2019

Editage

Paperpal

R Discovery

Mind the Graph

R Discovery Prime

R Discovery Prime

Analysis of Blocking and Scheduling for FPGA-Based Floating-Point Matrix Multiplication Analyse du blocage et de l’ordonnancement d’une multiplication matricielle à virgule flottante sur un FPGA

Abstract

Talk to us

Similar Papers

More From: Canadian Journal of Electrical and Computer Engineering