A Generalized Zero-Shot Quantization of Deep Convolutional Neural Networks Via Learned Weights Statistics

Prasen Kumar Sharma,Vikram Nelvoy Rajendiran,Arun Abraham

doi:10.1109/tmm.2021.3134158

Prasen Kumar Sharma, Vikram Nelvoy Rajendiran + Show 1 more

Open Access

https://doi.org/10.1109/tmm.2021.3134158

Copy DOI

Abstract

Quantizing the floating-point weights and activations of deep convolutional neural networks to fixed-point representation yields reduced memory footprints and inference time. Recently, efforts have been afoot towards zero-shot quantization that does not require original unlabelled training samples of a given task. These best-published works heavily rely on the learned batch normalization (BN) parameters to infer the range of the activations for quantization. In particular, these methods are built upon either empirical estimation framework or the data distillation approach, for computing the range of the activations. However, the performance of such schemes severely degrades when presented with a network that does not accommodate BN layers. In this line of thought, we propose a generalized zero-shot quantization (GZSQ) framework that neither requires original data nor relies on BN layer statistics. We have utilized the data distillation approach and leveraged only the pre-trained weights of the model to estimate enriched data for range calibration of the activations. To the best of our knowledge, this is the first work that utilizes the distribution of the pretrained weights to assist the process of zero-shot quantization. The proposed scheme has significantly outperformed the existing zero-shot works, e.g., an improvement of ~ 33% in classification accuracy for MobileNetV2 and several other models that are w & w/o BN layers, for a variety of tasks. We have also demonstrated the efficacy of the proposed work across multiple open-source quantization frameworks. Importantly, our work is the first attempt towards the post-training zero-shot quantization of futuristic unnormalized deep neural networks.

Talk to us

Join us for a 30 min session where you can share your feedback and ask us any queries you have

Schedule a call

Journal: IEEE Transactions on Multimedia	Publication Date: Jan 1, 2023
Citations: 5	License type: publisher-specific, author manuscript

R Discovery Prime

R Discovery Prime

A Generalized Zero-Shot Quantization of Deep Convolutional Neural Networks Via Learned Weights Statistics

Abstract

Talk to us

Similar Papers

More From: IEEE Transactions on Multimedia

Lead the way for us

Similar Papers

Supp1-3134158.pdf
Prasen Kumar Sharma
-
Prasen Kumar SharmaPrasen Kumar Sharma
10 Dec 2021
Supp1-3134158.pdf
Prasen Kumar Sharma

Training high-performance and large-scale deep neural networks with full 8-bit integers
Yukuan Yang ... Guoqi Li
Neural Networks | VOL. 125
Yukuan Yang, et. al.Yukuan Yang ... Guoqi Li
15 Jan 2020
Neural Networks | VOL. 125

Single-Bit-per-Weight Deep Convolutional Neural Networks without Batch-Normalization Layers for Embedded Systems
Mark D Mcdonnell ... Andrevan Schaik
-
Mark D Mcdonnell, et. al.Mark D Mcdonnell ... Andrevan Schaik
01 Jul 2019
01 Jul 2019

Continual Learning With Extended Kronecker-Factored Approximate Curvature
Janghyeon Lee ... Junmo Kim
-
Janghyeon Lee, et. al.Janghyeon Lee ... Junmo Kim
01 Jun 2020
01 Jun 2020

Editage

Paperpal

R Discovery

Mind the Graph

R Discovery Prime

R Discovery Prime

A Generalized Zero-Shot Quantization of Deep Convolutional Neural Networks Via Learned Weights Statistics

Abstract

Talk to us

Similar Papers

More From: IEEE Transactions on Multimedia