Accelerating Attention Mechanism on FPGAs based on Efficient Reconfigurable Systolic Array

Wenhua Ye,Joey Zhou,Kenli Li,Xu Zhou,Cen Chen

doi:10.1145/3549937

Abstract

Transformer model architectures have recently received great interest in natural language, machine translation, and computer vision, where attention mechanisms are their building blocks. However, the attention mechanism is expensive because of its intensive matrix computations and complicated data flow. The existing hardware architecture has some disadvantages for the computing structure of attention, such as inflexibility and low efficiency. Most of the existing papers accelerate attention by reducing the amount of computation through various pruning algorithms, which will affect the results to a certain extent with different sparsity. This paper proposes the hardware accelerator for the multi-head attention (MHA) on field-programmable gate arrays (FPGAs) with reconfigurable architecture, efficient systolic array, and hardware-friendly radix-2 softmax. We propose a novel method called Four inputs Processing Element (FPE) to double the computation rate of the data-aware systolic array (SA) and make it efficient and load balance. Especially, the computation framework is well designed to ensure the utilization of SA efficiently. Our design is evaluated on a Xilinx Alveo U250 card, and the proposed architecture achieves 51.3×, 17.3× improvement in latency, and 54.4×, 17.9× energy savings compared to CPU and GPU.

Talk to us

Join us for a 30 min session where you can share your feedback and ask us any queries you have

Schedule a call

R Discovery Prime

R Discovery Prime

Accelerating Attention Mechanism on FPGAs based on Efficient Reconfigurable Systolic Array

Abstract

Talk to us

Similar Papers

More From: ACM Transactions on Embedded Computing Systems

Lead the way for us

Journal: ACM Transactions on Embedded Computing Systems	Publication Date: Nov 9, 2023
Citations: 9

Similar Papers

Design for diagnosability and diagnostic strategies of WSI array architectures
Kuochen Wang ... Wang-Dauh Tseng
-
Kuochen Wang, et. al. Kuochen Wang ... Wang-Dauh Tseng
19 Jan 1994
19 Jan 1994

Prototyping time- and space-efficient computations of algebraic operations over dynamically reconfigurable systems modeled by rewriting-logic
M Ayala-Rincón ... R P Jacobi
ACM Transactions on Design Automation of Electronic Systems | VOL. 11
M Ayala-Rincón, et. al.M Ayala-Rincón ... R P Jacobi
01 Apr 2006
ACM Transactions on Design Automation of Electronic Systems | VOL. 11

Maestro: A Memory-on-Logic Architecture for Coordinated Parallel Use of Many Systolic Arrays
H. T. Kung ... Bradley McDanel
-
H. T. Kung, et. al.H. T. Kung ... Bradley McDanel
01 Jul 2019
01 Jul 2019

Wafer-scale integration and two-level pipelined implementations of systolic arrays
H.T Kung ... Monica S Lam
Journal of Parallel and Distributed Computing | VOL. 1
H.T Kung, et. al.H.T Kung ... Monica S Lam
01 Aug 1984
Journal of Parallel and Distributed Computing | VOL. 1

Editage

Paperpal

R Discovery

Mind the Graph

R Discovery Prime

R Discovery Prime

Accelerating Attention Mechanism on FPGAs based on Efficient Reconfigurable Systolic Array

Abstract

Talk to us

Similar Papers

More From: ACM Transactions on Embedded Computing Systems