Spatial- and time- division multiplexing in CNN accelerator

Tetsuro Nakamura,Shogo Saito,Kei Fujimoto,Masashi Kaneko,Akinori Shiraga

doi:10.1016/j.parco.2022.102922

Abstract

With the widespread use of real-time data analysis by artificial intelligence (AI), the integration of accelerators is attracting attention from the perspectives of their low power consumption and low latency. The objective of this research is to increase accelerator resource efficiency and further reduce power consumption by sharing accelerators among multiple users while maintaining real-time performance. To achieve the accelerator-sharing system, we define three requirements: high device utilization, fair device utilization among users, and real-time performance. Targeting the AI inference use case, this paper proposes a system that shares a field-programmable gate array (FPGA) among multiple users by switching the convolutional neural network (CNN) models stored in the device memory on the FPGA, while satisfying the three requirements. The proposed system uses different behavioral models for workloads with predictable and unpredictable data arrival timing. For the workloads with predictable data arrival timing, the system uses spatial-division multiplexing of the FPGA device memory to achieve real-time performance and high device utilization. Specifically, the FPGA device memory controller of the system transparently preloads and caches the CNN models into the FPGA device memory before the data arrival. For workloads with unpredictable data arrival timing, the system transfers CNN models to the FPGA device memory upon data arrival using time-division multiplexing of FPGA device memory. In the latter case of unpredictable workloads, the switch cost between CNN models is non-negligible to achieve real-time performance and high device utilization, so the system integrates a new scheduling algorithm that considers the switch time of the CNN models. For both predictable and unpredictable workloads, user fairness is achieved by using an ageing technique in the scheduling algorithm that increases the priority of jobs in accordance with the job waiting time. The evaluation results show that the scheduling overhead of the proposed system is negligible for both predictable and unpredictable workloads providing practical real-time performance. For unpredictable workloads, the new scheduling algorithm improves fairness by 24%–94% and resource efficiency by 31%–33% compared to traditional algorithms using first-come first-served or round-robin. For predictable workloads, the system improves fairness by 50.5 % compared to first-come first-served and achieves 99.5 % resource efficiency.

Full Text

Published Version

View

Talk to us

Join us for a 30 min session where you can share your feedback and ask us any queries you have

Schedule a call

Journal: Parallel Computing	Publication Date: Mar 24, 2022
Citations: 1	License type: cc-by-nc-nd

R Discovery Prime

Spatial- and time- division multiplexing in CNN accelerator

Abstract

Published Version

Talk to us

Similar Papers

More From: Parallel Computing

Lead the way for us

Similar Papers

Time-Division Multiplexing for FPGA Considering CNN Model Switch Time
Tetsuro Nakamura ... Shogo Saito
-
Tetsuro Nakamura, et. al.Tetsuro Nakamura ... Shogo Saito
01 Jun 2021
01 Jun 2021

Artificial intelligence: finding the intersection of predictive modeling and clinical utility
Karthik Ravi
Gastrointestinal Endoscopy | VOL. 93
Karthik RaviKarthik Ravi
07 Mar 2021
Gastrointestinal Endoscopy | VOL. 93

An FPGA Design Framework for CNN Sparsification and Acceleration
Sicheng Li ... Yu Wang
-
Sicheng Li, et. al.Sicheng Li ... Yu Wang
01 Apr 2017
01 Apr 2017

Evaluating a Comparing Deep Learning Architectures for Blood Glucose Prediction
Touria El Idrissi ... Ali Idri
-
Touria El Idrissi, et. al.Touria El Idrissi ... Ali Idri
01 Jan 2020
01 Jan 2020

Editage

Paperpal

R Discovery

Mind the Graph

R Discovery Prime

Spatial- and time- division multiplexing in CNN accelerator

Abstract

Published Version

Talk to us

Similar Papers

More From: Parallel Computing