Power-Optimal Mapping of CNN Applications to Cloud-Based Multi-FPGA Platforms

Junnan Shan,Mario R Casu,Mihai T Lazarescu,Luciano Lavagno,Jordi Cortadella

doi:10.1109/tcsii.2020.2998284

Abstract

Multi-FPGA platforms like Amazon Web Services F1 are perfect to accelerate multi-kernel pipelined applications, like Convolutional Neural Networks (CNNs). To reduce energy consumption, we propose to upload at runtime the best power-optimized CNN implementation for a given throughput constraint. Our design method gives the best number of parallel instances of each kernel, their allocation to the FPGAs, the number of powered-on FPGAs and their clock frequency. This is obtained by solving a mixed-integer, non-linear optimization problem that models power and performance of each component, as well as the duration of the computation phases—data transfer between a host CPU and the FPGA memory (typically DDR), data transfer between DDR and FPGA, and FPGA computation. The results show that the power saved compared to simply clock gating the fastest implementation is obviously very high, but it is also much more significant than simply scaling the frequency of the fastest implementation or replicating the slowest implementation on multiple FPGAs.

Talk to us

Join us for a 30 min session where you can share your feedback and ask us any queries you have

Schedule a call

Journal: IEEE Transactions on Circuits and Systems II: Express Briefs	Publication Date: Dec 1, 2020
Citations: 16	License type: other-oa

R Discovery Prime

R Discovery Prime

Power-Optimal Mapping of CNN Applications to Cloud-Based Multi-FPGA Platforms

Abstract

Talk to us

Similar Papers

More From: IEEE Transactions on Circuits and Systems II: Express Briefs

Lead the way for us

Similar Papers

CNN-on-AWS: Efficient Allocation of Multikernel Applications on Multi-FPGA Platforms
Junnan Shan ... Jordi Cortadella
IEEE Transactions on Computer-Aided Design of Integrated Circuits and Systems | VOL. 40
Junnan Shan, et. al.Junnan Shan ... Jordi Cortadella
15 May 2020
IEEE Transactions on Computer-Aided Design of Integrated Circuits and Systems | VOL. 40

Towards Efficient Neural Network Model Parallelism on Multi-FPGA Platforms
David Rodríguez Agut ... Rafael Tornero
-
David Rodríguez Agut, et. al.David Rodríguez Agut ... Rafael Tornero
01 Apr 2023
01 Apr 2023

Novel Design partitioning technique for ASIC prototyping on multi-FPGA platforms using Graph Deep Learning
Divyasree Tummalapalli ... Chiranjeevi Kunapareddy
-
Divyasree Tummalapalli, et. al.Divyasree Tummalapalli ... Chiranjeevi Kunapareddy
24 Oct 2022
24 Oct 2022

NnDPI: A Novel Deep Packet Inspection Technique Using Word Embedding, Convolutional and Recurrent Neural Networks
Mahmoud Bahaa ... Khaled Adel
-
Mahmoud Bahaa, et. al.Mahmoud Bahaa ... Khaled Adel
24 Oct 2020
24 Oct 2020

Editage

Paperpal

R Discovery

Mind the Graph

R Discovery Prime

R Discovery Prime

Power-Optimal Mapping of CNN Applications to Cloud-Based Multi-FPGA Platforms

Abstract

Talk to us

Similar Papers

More From: IEEE Transactions on Circuits and Systems II: Express Briefs