Simultaneous Multithreaded Matrix Processor

Mostafa I Soliman,Elsayed A Elsayed

doi:10.1142/s0218126615501145

Abstract

This paper proposes a simultaneous multithreaded matrix processor (SMMP) to improve the performance of data-parallel applications by exploiting instruction-level parallelism (ILP) data-level parallelism (DLP) and thread-level parallelism (TLP). In SMMP, the well-known five-stage pipeline (baseline scalar processor) is extended to execute multi-scalar/vector/matrix instructions on unified parallel execution datapaths. SMMP can issue four scalar instructions from two threads each cycle or four vector/matrix operations from one thread, where the execution of vector/matrix instructions in threads is done in round-robin fashion. Moreover, this paper presents the implementation of our proposed SMMP using VHDL targeting FPGA Virtex-6. In addition, the performance of SMMP is evaluated on some kernels from the basic linear algebra subprograms (BLAS). Our results show that, the hardware complexity of SMMP is 5.68 times higher than the baseline scalar processor. However, speedups of 4.9, 6.09, 6.98, 8.2, 8.25, 8.72, 9.36, 11.84 and 21.57 are achieved on BLAS kernels of applying Givens rotation, scalar times vector plus another, vector addition, vector scaling, setting up Givens rotation, dot-product, matrix–vector multiplication, Euclidean length, and matrix–matrix multiplications, respectively. The average speedup over the baseline is 9.55 and the average speedup over complexity is 1.68. Comparing with Xilinx MicroBlaze, the complexity of SMMP is 6.36 times higher, however, its speedup ranges from 6.87 to 12.07 on vector/matrix kernels, which is 9.46 in average.

Talk to us

Join us for a 30 min session where you can share your feedback and ask us any queries you have

Schedule a call

R Discovery Prime

R Discovery Prime

Simultaneous Multithreaded Matrix Processor

Abstract

Talk to us

Similar Papers

More From: Journal of Circuits, Systems and Computers

Lead the way for us

Similar Papers

FPGA implementation and performance evaluation of a simultaneous multithreaded matrix processor
Mostafa I Soliman ... Elsayed A Elsayed
-
Mostafa I Soliman, et. al.Mostafa I Soliman ... Elsayed A Elsayed
01 Dec 2014
01 Dec 2014

Converting thread-level parallelism to instruction-level parallelism via simultaneous multithreading
Jack L Lo ... Joel S Emer
ACM Transactions on Computer Systems | VOL. 15
Jack L Lo, et. al.Jack L Lo ... Joel S Emer
01 Aug 1997
ACM Transactions on Computer Systems | VOL. 15

Simultaneous multithreading-based routers
K Vibhatavanij ... Nian-Feng Tzeng
-
K Vibhatavanij, et. al.K Vibhatavanij ... Nian-Feng Tzeng
21 Aug 2000
21 Aug 2000

The Case for Speculative Multithreading on SMT Processors
Haitham Akkary ... Sébastien Hily
-
Haitham Akkary, et. al.Haitham Akkary ... Sébastien Hily
01 Jan 1999
01 Jan 1999

Editage

Paperpal

R Discovery

Mind the Graph

R Discovery Prime

R Discovery Prime

Simultaneous Multithreaded Matrix Processor

Abstract

Talk to us

Similar Papers

More From: Journal of Circuits, Systems and Computers