Block Convolution: Toward Memory-Efficient Inference of Large-Scale CNNs on FPGA

Gang Li,Jian Cheng,Fanrong Li,Zejian Liu

doi:10.1109/tcad.2021.3082868

Abstract

Deep convolutional neural networks have achieved remarkable progress in recent years. However, the large volume of intermediate results generated during inference poses a significant challenge to the accelerator design for resource-constrained field-programmable gate array (FPGA). Due to the limited on-chip storage, partial results of intermediate layers are frequently transferred back and forth between on-chip memory and off-chip DRAM, leading to a nonnegligible increase in latency and energy consumption. In this article, we propose block convolution, a hardware-friendly, simple, yet efficient convolution operation that can completely avoid the off-chip transfer of intermediate feature maps at runtime. The fundamental idea of block convolution is to eliminate the dependency of feature map tiles in the spatial dimension when spatial tiling is used, which is realized by splitting a feature map into independent blocks so that convolution can be performed separately on individual blocks. We conduct extensive experiments to demonstrate the efficacy of the proposed block convolution on both the algorithm side and the hardware side. Specifically, we evaluate block convolution on: 1) VGG-16, ResNet-18, ResNet-50, and MobileNet-V1 for the ImageNet classification task; 2) SSD and FPN for the COCO object detection task; and 3) VDSR for the Set5 single-image superresolution task. Experimental results demonstrate that comparable or higher accuracy can be achieved with block convolution. We also showcase two CNN accelerators via algorithm/hardware co-design based on block convolution on memory-limited FPGAs, and evaluation shows that both accelerators substantially outperform the baseline without off-chip transfer of intermediate feature maps.

Full Text

Published version (

Free)

Open DOI Link

Talk to us

Join us for a 30 min session where you can share your feedback and ask us any queries you have

Schedule a call

R Discovery Prime

R Discovery Prime

Block Convolution: Toward Memory-Efficient Inference of Large-Scale CNNs on FPGA

Abstract

Talk to us

Similar Papers

More From: IEEE Transactions on Computer-Aided Design of Integrated Circuits and Systems

Lead the way for us

Journal: IEEE Transactions on Computer-Aided Design of Integrated Circuits and Systems	Publication Date: May 22, 2021
Citations: 22

Similar Papers

A Block-Based and Highly Parallel CNN Accelerator for Seed Sorting
Xiaoting Sang ... Yang Li
Journal of Electrical and Computer Engineering | VOL. 2022
Xiaoting Sang, et. al.Xiaoting Sang ... Yang Li
17 Nov 2022
Journal of Electrical and Computer Engineering | VOL. 2022

Block convolution: Towards memory-efficient inference of large-scale CNNs on FPGA
Gang Li ... Tianli Zhao
-
Gang Li, et. al.Gang Li ... Tianli Zhao
01 Mar 2018
01 Mar 2018

Evolutionary bin packing for memory-efficient dataflow inference acceleration on FPGA
Mairin Kroes ... Sorin Cotofana
-
Mairin Kroes, et. al.Mairin Kroes ... Sorin Cotofana
25 Jun 2020
25 Jun 2020

Caching Hybrid Rotation: A Memory Access Optimization Method for CNN on FPGA
Dong Dong ... Xuekai Wei
Journal of Circuits, Systems and Computers | VOL. 32
Dong Dong, et. al.Dong Dong ... Xuekai Wei
04 Mar 2023
Journal of Circuits, Systems and Computers | VOL. 32

Editage

Paperpal

R Discovery

Mind the Graph

R Discovery Prime

R Discovery Prime

Block Convolution: Toward Memory-Efficient Inference of Large-Scale CNNs on FPGA

Abstract

Talk to us

Similar Papers

More From: IEEE Transactions on Computer-Aided Design of Integrated Circuits and Systems