Optimizing Recurrent Neural Networks: A Study on Gradient Normalization of Weights for Enhanced Training Efficiency

Xinyi Wu,Bingjie Xiang,Xingwang Huang,Chaopeng Li,Huaizheng Lu,Weifang Huang

doi:10.3390/app14156578

Abstract

Recurrent Neural Networks (RNNs) are classical models for processing sequential data, demonstrating excellent performance in tasks such as natural language processing and time series prediction. However, during the training of RNNs, the issues of vanishing and exploding gradients often arise, significantly impacting the model’s performance and efficiency. In this paper, we investigate why RNNs are more prone to gradient problems compared to other common sequential networks. To address this issue and enhance network performance, we propose a method for gradient normalization of network weights. This method suppresses the occurrence of gradient problems by altering the statistical properties of RNN weights, thereby improving training effectiveness. Additionally, we analyze the impact of weight gradient normalization on the probability-distribution characteristics of model weights and validate the sensitivity of this method to hyperparameters such as learning rate. The experimental results demonstrate that gradient normalization enhances the stability of model training and reduces the frequency of gradient issues. On the Penn Treebank dataset, this method achieves a perplexity level of 110.89, representing an 11.48% improvement over conventional gradient descent methods. For prediction lengths of 24 and 96 on the ETTm1 dataset, Mean Absolute Error (MAE) values of 0.778 and 0.592 are attained, respectively, resulting in 3.00% and 6.77% improvement over conventional gradient descent methods. Moreover, selected subsets of the UCR dataset show an increase in accuracy ranging from 0.4% to 6.0%. The gradient normalization method enhances the ability of RNNs to learn from sequential and causal data, thereby holding significant implications for optimizing the training effectiveness of RNN-based models.

Talk to us

Join us for a 30 min session where you can share your feedback and ask us any queries you have

Schedule a call

R Discovery Prime

R Discovery Prime

Optimizing Recurrent Neural Networks: A Study on Gradient Normalization of Weights for Enhanced Training Efficiency

Abstract

Talk to us

Similar Papers

More From: Applied Sciences

Lead the way for us

Journal: Applied Sciences	Publication Date: Jul 27, 2024
License type: CC BY 4.0

Similar Papers

Wavefront parallelization of recurrent neural networks on multi-core architectures
Robin Kumar Sharma ... Marc Casas
-
Robin Kumar Sharma, et. al.Robin Kumar Sharma ... Marc Casas
29 Jun 2020
29 Jun 2020

Constructive training of recurrent neural networks using hybrid optimization
Niranjan Subrahmanya ... Yung C Shin
Neurocomputing | VOL. 73
Niranjan Subrahmanya, et. al.Niranjan Subrahmanya ... Yung C Shin
03 Jul 2010
Neurocomputing | VOL. 73

Improved recurrent NARX neural network model for state of charge estimation of lithium-ion battery using pso algorithm
M S Hossain Lipu ... A Ayob
-
M S Hossain Lipu, et. al.M S Hossain Lipu ... A Ayob
01 Apr 2018
01 Apr 2018

Sequence-discriminative training of recurrent neural networks
Paul Voigtlaender ... Ralf Schluter
-
Paul Voigtlaender, et. al.Paul Voigtlaender ... Ralf Schluter
01 Apr 2015
01 Apr 2015

Editage

Paperpal

R Discovery

Mind the Graph

R Discovery Prime

R Discovery Prime

Optimizing Recurrent Neural Networks: A Study on Gradient Normalization of Weights for Enhanced Training Efficiency

Abstract

Talk to us

Similar Papers

More From: Applied Sciences