Mixed-precision quantized neural networks with progressively decreasing bitwidth

Tianshu Chu,Qin Luo,Jie Yang,Xiaolin Huang

doi:10.1016/j.patcog.2020.107647

Abstract

Efficient model inference is an important and practical issue in the deployment of deep neural networks on resource constraint platforms. Network quantization addresses this problem effectively by leveraging low-bit representation and arithmetic that could be conducted on dedicated embedded systems. In the previous works, the parameter bitwidth is set homogeneously and there is a trade-off between superior performance and aggressive compression. Actually, the stacked network layers, which are generally regarded as hierarchical feature extractors, contribute diversely to the overall performance. For a well-trained neural network, the feature distributions of different categories are organized gradually as the network propagates forward. Hence the capability requirement on the subsequent feature extractors is reduced. It indicates that the neurons in posterior layers could be assigned with lower bitwidth for quantized neural networks. Based on this observation, a simple yet effective mixed-precision quantized neural network with progressively decreasing bitwidth is proposed to improve the trade-off between accuracy and compression. Extensive experiments on typical network architectures and benchmark datasets demonstrate that the proposed method could achieve better or comparable results while reducing the memory space for quantized parameters by more than 25% in comparison with the homogeneous counterparts. In addition, the results also demonstrate that the higher-precision bottom layers could boost the 1-bit network performance appreciably due to a better preservation of the original image information while the lower-precision posterior layers contribute to the regularization of k−bit networks.

Talk to us

Join us for a 30 min session where you can share your feedback and ask us any queries you have

Schedule a call

R Discovery Prime

R Discovery Prime

Mixed-precision quantized neural networks with progressively decreasing bitwidth

Abstract

Talk to us

Similar Papers

More From: Pattern Recognition

Lead the way for us

Journal: Pattern Recognition	Publication Date: Sep 24, 2020
Citations: 27

Similar Papers

An Ordered Aggregation-Based Ensemble Selection Method of Lightweight Deep Neural Networks With Random Initialization
Lin He ... Lijun Peng
IEEE Access | VOL. 10
Lin He, et. al.Lin He ... Lijun Peng
01 Jan 2021
IEEE Access | VOL. 10

DORY: Automatic End-to-End Deployment of Real-World DNNs on Low-Cost IoT MCUs
Alessio Burrello ... Nazareno Bruschi
IEEE Transactions on Computers | VOL. 70
Alessio Burrello, et. al.Alessio Burrello ... Nazareno Bruschi
01 Aug 2021
IEEE Transactions on Computers | VOL. 70

A convolutional neural network for traffic information sensing from social media text
Yuanyuan Chen ... Yisheng Lv
-
Yuanyuan Chen, et. al.Yuanyuan Chen ... Yisheng Lv
01 Oct 2017
01 Oct 2017

Deep learning acceleration on edge devices with algorithm/hardware co-design
Mengshu Sun
-
Mengshu SunMengshu Sun
10 Feb 2023
10 Feb 2023

Editage

Paperpal

R Discovery

Mind the Graph

R Discovery Prime

R Discovery Prime

Mixed-precision quantized neural networks with progressively decreasing bitwidth

Abstract

Talk to us

Similar Papers

More From: Pattern Recognition