Profiling the real world potential of neural network compression

Joe Lorentz,Djamila Aouada,Thomas Hartmann,Assaad Moawad

doi:10.1109/coins54846.2022.9854973

Abstract

Many real world computer vision applications are required to run on hardware with limited computing power, often referred to as "edge devices". The state of the art in computer vision continues towards ever bigger and deeper neural networks with equally rising computational requirements. Model compression methods promise to substantially reduce the computation time and memory demands with little to no impact on the model robustness. However, evaluation of the compression is mostly based on theoretic speedups in terms of required floating-point operations. This work offers a tool to profile the actual speedup offered by several compression algorithms. Our results show a significant discrepancy between the theoretical and actual speedup on various hardware setups. Furthermore, we show the potential of model compressions and highlight the importance of selecting the right compression algorithm for a target task and hardware. The code to reproduce our experiments is available at https://hub.datathings.com/papers/2022-coins.

Full Text