Model Parallelism Optimization for CNN FPGA Accelerator

Jinnan Wang,Xiaoli Zhi,Weiqin Tong

doi:10.3390/a16020110

Abstract

Convolutional neural networks (CNNs) have made impressive achievements in image classification and object detection. For hardware with limited resources, it is not easy to achieve CNN inference with a large number of parameters without external storage. Model parallelism is an effective way to reduce resource usage by distributing CNN inference among several devices. However, parallelizing a CNN model is not easy, because CNN models have an essentially tightly-coupled structure. In this work, we propose a novel model parallelism method to decouple the CNN structure with group convolution and a new channel shuffle procedure. Our method could eliminate inter-device synchronization while reducing the memory footprint of each device. Using the proposed model parallelism method, we designed a parallel FPGA accelerator for the classic CNN model ShuffleNet. This accelerator was further optimized with features such as aggregate read and kernel vectorization to fully exploit the hardware-level parallelism of the FPGA. We conducted experiments with ShuffleNet on two FPGA boards, each of which had an Intel Arria 10 GX1150 and 16GB DDR3 memory. The experimental results showed that when using two devices, ShuffleNet achieved a 1.42× speed increase and reduced its memory footprint by 34%, as compared to its non-parallel counterpart, while maintaining accuracy.

Talk to us

Join us for a 30 min session where you can share your feedback and ask us any queries you have

Schedule a call

Journal: Algorithms	Publication Date: Feb 14, 2023
Citations: 7	License type: CC BY 4.0

R Discovery Prime

R Discovery Prime

Model Parallelism Optimization for CNN FPGA Accelerator

Abstract

Talk to us

Similar Papers

More From: Algorithms

Lead the way for us

Similar Papers

Arithmetic Coding-Based 5-Bit Weight Encoding and Hardware Decoder for CNN Inference in Edge Devices
Jong Hun Lee ... Joonho Kong
IEEE Access | VOL. 9
Jong Hun Lee, et. al.Jong Hun Lee ... Joonho Kong
01 Jan 2020
IEEE Access | VOL. 9

An Efficient Design Flow for Accelerating Complicated-connected CNNs on a Multi-FPGA Platform
Deguang Wang ... Junzhong Shen
-
Deguang Wang, et. al.Deguang Wang ... Junzhong Shen
05 Aug 2019
05 Aug 2019

A Depthwise Separable Convolution Architecture for CNN Accelerator
Harsh Srivastava ... Kishor Sarawadekar
-
Harsh Srivastava, et. al.Harsh Srivastava ... Kishor Sarawadekar
07 Oct 2020
07 Oct 2020

CPU-Accelerator Co-Scheduling for CNN Acceleration at the Edge
Yeongmin Kim ... Joonho Kong
IEEE Access | VOL. 8
Yeongmin Kim, et. al.Yeongmin Kim ... Joonho Kong
01 Jan 2020
IEEE Access | VOL. 8

Editage

Paperpal

R Discovery

Mind the Graph

R Discovery Prime

R Discovery Prime

Model Parallelism Optimization for CNN FPGA Accelerator

Abstract

Talk to us

Similar Papers

More From: Algorithms