An FPGA-Based Reconfigurable Accelerator for Low-Bit DNN Training

Haikuo Shao,Jinming Lu,Jun Lin,Zhongfeng Wang

doi:10.1109/isvlsi51109.2021.00054

Abstract

In recent years, deep neural networks (DNNs) have shown outstanding performance in various tasks. Training DNNs on resource-constrained edge platforms is required for online learning and privacy concerns. However, the training process of DNNs requires enormous computation and memory resources. Therefore, the hardware implementations of the training process with low-precision integer arithmetic are attracting extensive attention, considering their advantages on computation, storage, and energy consumption. In this paper, we propose an FPGA-based reconfigurable accelerator for DNN training with full 8-bit integer arithmetic. First, a reconfigurable processing element in a unified architecture is designed, which is flexible to support various computation patterns during the training process. Second, a two-stage scaling and rounding scheme is introduced to scale intermediate results to low-bit data for minimal memory usage, while retaining data accuracy to the maximum extent. Finally, based on the widely used softmax classification function, an optimized architecture is developed to calculate the cross-entropy loss function on-chip to initiate the backward propagation. Experimental results show that our design can reach 771 GOPS and 47.38 GOPS/W in terms of performance and energy efficiency, respectively. The comparison results demonstrate that our work significantly outperforms prior works.

Talk to us

Join us for a 30 min session where you can share your feedback and ask us any queries you have

Schedule a call

R Discovery Prime

R Discovery Prime

An FPGA-Based Reconfigurable Accelerator for Low-Bit DNN Training

Abstract

Talk to us

Similar Papers

Lead the way for us

Similar Papers

A Reconfigurable DNN Training Accelerator on FPGA
Jinming Lu ... Jun Lin
-
Jinming Lu, et. al.Jinming Lu ... Jun Lin
24 Sep 2020
24 Sep 2020

A Framework for Distributed Deep Neural Network Training with Heterogeneous Computing Platforms
Bontak Gu ... Arslan Munir
-
Bontak Gu, et. al.Bontak Gu ... Arslan Munir
01 Dec 2019
01 Dec 2019

Neuroevolution in Deep Neural Networks: Current Trends and Future Challenges
Edgar Galvan ... Peter Mooney
IEEE Transactions on Artificial Intelligence | VOL. 2
Edgar Galvan, et. al.Edgar Galvan ... Peter Mooney
04 May 2021
IEEE Transactions on Artificial Intelligence | VOL. 2

A convergence analysis of Nesterov’s accelerated gradient method in training deep linear neural networks
Xin Liu ... Zhisong Pan
Information Sciences | VOL. 612
Xin Liu, et. al.Xin Liu ... Zhisong Pan
05 Sep 2022
Information Sciences | VOL. 612

Editage

Paperpal

R Discovery

Mind the Graph

R Discovery Prime

R Discovery Prime

An FPGA-Based Reconfigurable Accelerator for Low-Bit DNN Training

Abstract

Talk to us

Similar Papers