Investigations on hessian-free optimization for cross-entropy training of deep neural networks

Simon Wiesler,Jian Xue,Jinyu Li

doi:10.21437/interspeech.2013-734

Simon Wiesler, Jian Xue + Show 1 more

Open Access

https://doi.org/10.21437/interspeech.2013-734

Copy DOI

Abstract

Context-dependent deep neural network HMMs have been shown to achieve recognition accuracy superior to Gaussian mixture models in a number of recent works. Typically, neural networks are optimized with stochastic gradient descent. On large datasets, stochastic gradient descent improves quickly during the beginning of the optimization. But since it does not make use of second order information, its asymptotic convergence behavior is slow. In regions with pathological curvature, stochastic gradient descent may almost stagnate and thereby falsely indicate convergence. Another drawback of stochastic gradient descent is that it can only be parallelized within minibatches. The Hessian-free algorithm is a second order batch optimization algorithm that does not suffer from these problems. In a recent work, Hessian-free optimization has been applied to a training of deep neural networks according to a sequence criterion. In that work, improvements in accuracy and training time have been reported. In this paper, we analyze the properties of the Hessian-free optimization algorithm and investigate whether it is suited for cross-entropy training of deep neural networks as well.

Talk to us

Join us for a 30 min session where you can share your feedback and ask us any queries you have

Schedule a call

R Discovery Prime

R Discovery Prime

Investigations on hessian-free optimization for cross-entropy training of deep neural networks

Abstract

Talk to us

Similar Papers

Lead the way for us

Publication Date: Aug 25, 2013
Citations: 20	License type: other-oa

Similar Papers

Investigation of stochastic Hessian-Free optimization in Deep neural networks for speech recognition
Zhao You ... Bo Xu
-
Zhao You, et. al.Zhao You ... Bo Xu
01 Sep 2014
01 Sep 2014

Neuroevolution in Deep Neural Networks: Current Trends and Future Challenges
Edgar Galvan ... Peter Mooney
IEEE Transactions on Artificial Intelligence | VOL. 2
Edgar Galvan, et. al.Edgar Galvan ... Peter Mooney
04 May 2021
IEEE Transactions on Artificial Intelligence | VOL. 2

Non-convergence of stochastic gradient descent in the training of deep neural networks
Patrick Cheridito ... Florian Rossmannek
Journal of Complexity | VOL. 64
Patrick Cheridito, et. al.Patrick Cheridito ... Florian Rossmannek
27 Nov 2020
Journal of Complexity | VOL. 64

A Framework for Distributed Deep Neural Network Training with Heterogeneous Computing Platforms
Bontak Gu ... Arslan Munir
-
Bontak Gu, et. al.Bontak Gu ... Arslan Munir
01 Dec 2019
01 Dec 2019

Editage

Paperpal

R Discovery

Mind the Graph

R Discovery Prime

R Discovery Prime

Investigations on hessian-free optimization for cross-entropy training of deep neural networks

Abstract

Talk to us

Similar Papers