Embedding Principle: A Hierarchical Structure of Loss Landscape of Deep Neural Networks

Yaoyu Zhang,Tao Luo Null,Yuqing Li,Zhi-Qin John Xu,Zhongwang Zhang

doi:10.4208/jml.220108

Abstract

We prove a general Embedding Principle of loss landscape of deep neural networks (NNs) that unravels a hierarchical structure of the loss landscape of NNs, i.e., loss landscape of an NN contains all critical points of all the narrower NNs. This result is obtained by constructing a class of critical embeddings which map any critical point of a narrower NN to a critical point of the target NN with the same output function. By discovering a wide class of general compatible critical embeddings, we provide a gross estimate of the dimension of critical submanifolds embedded from critical points of narrower NNs. We further prove an irreversiblility property of any critical embedding that the number of negative/zero/positive eigenvalues of the Hessian matrix of a critical point may increase but never decrease as an NN becomes wider through the embedding. Using a special realization of general compatible critical embedding, we prove a stringent necessary condition for being a "truly-bad" critical point that never becomes a strict-saddle point through any critical embedding. This result implies the commonplace of strict-saddle points in wide NNs, which may be an important reason underlying the easy optimization of wide NNs widely observed in practice.

Talk to us

Join us for a 30 min session where you can share your feedback and ask us any queries you have

Schedule a call

R Discovery Prime

R Discovery Prime

Embedding Principle: A Hierarchical Structure of Loss Landscape of Deep Neural Networks

Abstract

Talk to us

Similar Papers

More From: Journal of Machine Learning

Lead the way for us

Similar Papers

Comparing general and specialized word embeddings for biomedical named entity recognition.
Rigo E Ramos-Vargas ... Sulema Torres-Ramos
PeerJ Computer Science | VOL. 7
Rigo E Ramos-Vargas, et. al.Rigo E Ramos-Vargas ... Sulema Torres-Ramos
18 Feb 2021
PeerJ Computer Science | VOL. 7

Construction of Neural Networks that Do Not Have Critical Points Based on Hierarchical Structure
Tohru Nitta
International Journal of Advanced Computer Science and Applications | VOL. 4
Tohru NittaTohru Nitta
01 Jan 2013
International Journal of Advanced Computer Science and Applications | VOL. 4

Critical exponents from AdS/CFT with flavor
Andreas Karch ... Laurence G Yaffe
Journal of High Energy Physics | VOL. 2009
Andreas Karch, et. al.Andreas Karch ... Laurence G Yaffe
07 Sep 2009
Journal of High Energy Physics | VOL. 2009

Quantum thetas on noncommutative with general embeddings
Ee Chang-Young ... Hoil Kim
Journal of Physics A: Mathematical and Theoretical | VOL. 41
Ee Chang-Young, et. al.Ee Chang-Young ... Hoil Kim
26 Feb 2008
Journal of Physics A: Mathematical and Theoretical | VOL. 41

Editage

Paperpal

R Discovery

Mind the Graph

R Discovery Prime

R Discovery Prime

Embedding Principle: A Hierarchical Structure of Loss Landscape of Deep Neural Networks

Abstract

Talk to us

Similar Papers

More From: Journal of Machine Learning