Coordinate Descent on the Orthogonal Group for Recurrent Neural Network Training

Estelle Massart,Vinayak Abrol

doi:10.1609/aaai.v36i7.20742

Abstract

We address the poor scalability of learning algorithms for orthogonal recurrent neural networks via the use of stochastic coordinate descent on the orthogonal group, leading to a cost per iteration that increases linearly with the number of recurrent states. This contrasts with the cubic dependency of typical feasible algorithms such as stochastic Riemannian gradient descent, which prohibits the use of big network architectures. Coordinate descent rotates successively two columns of the recurrent matrix. When the coordinate (i.e., indices of rotated columns) is selected uniformly at random at each iteration, we prove convergence of the algorithm under standard assumptions on the loss function, stepsize and minibatch noise. In addition, we numerically show that the Riemannian gradient has an approximately sparse structure. Leveraging this observation, we propose a variant of our proposed algorithm that relies on the Gauss-Southwell coordinate selection rule. Experiments on a benchmark recurrent neural network training problem show that the proposed approach is a very promising step towards the training of orthogonal recurrent neural networks with big architectures.

Talk to us

Join us for a 30 min session where you can share your feedback and ask us any queries you have

Schedule a call

R Discovery Prime

R Discovery Prime

Coordinate Descent on the Orthogonal Group for Recurrent Neural Network Training

Abstract

Talk to us

Similar Papers

More From: Proceedings of the AAAI Conference on Artificial Intelligence

Lead the way for us

Journal: Proceedings of the AAAI Conference on Artificial Intelligence	Publication Date: Jun 28, 2022
Citations: 3

Similar Papers

Reinforced two-step-ahead weight adjustment technique for online training of recurrent neural networks.
Li-Chiu Chang ... Pin-An Chen
IEEE Transactions on Neural Networks and Learning Systems | VOL. 23
Li-Chiu Chang, et. al. Li-Chiu Chang ... Pin-An Chen
01 Aug 2012
IEEE Transactions on Neural Networks and Learning Systems | VOL. 23

A real-coded genetic algorithm for training recurrent neural networks
A Blanco ... M.C Pegalajar
Neural Networks | VOL. 14
A Blanco, et. al.A Blanco ... M.C Pegalajar
01 Jan 2001
Neural Networks | VOL. 14

New second-order algorithms for recurrent neural networks based on conjugate gradient
P Campolucci ... M Simonetti
-
P Campolucci, et. al.P Campolucci ... M Simonetti
04 May 1998
04 May 1998

A robust recurrent simultaneous perturbation stochastic approximation training algorithm for recurrent neural networks
Zhao Xu ... Danwei Wang
Neural Computing and Applications | VOL. 24
Zhao Xu, et. al.Zhao Xu ... Danwei Wang
14 Jun 2013
Neural Computing and Applications | VOL. 24

Editage

Paperpal

R Discovery

Mind the Graph

R Discovery Prime

R Discovery Prime

Coordinate Descent on the Orthogonal Group for Recurrent Neural Network Training

Abstract

Talk to us

Similar Papers

More From: Proceedings of the AAAI Conference on Artificial Intelligence