Sign-Based Gradient Descent With Heterogeneous Data: Convergence and Byzantine Resilience.

Richeng Jin,Xiaofan He,Tianfu Wu,Huaiyu Dai,Yufan Huang,Yuding Liu

doi:10.1109/tnnls.2023.3345367

Abstract

Communication overhead has become one of the major bottlenecks in the distributed training of modern deep neural networks. With such consideration, various quantization-based stochastic gradient descent (SGD) solvers have been proposed and widely adopted, among which signSGD with majority vote shows a promising direction because of its communication efficiency and robustness against Byzantine attackers. However, signSGD fails to converge in the presence of data heterogeneity, which is commonly observed in the emerging federated learning (FL) paradigm. In this article, a sufficient condition for the convergence of the sign-based gradient descent method is derived, based on which a novel magnitude-driven stochastic-sign-based gradient compressor is proposed to address the non-convergence issue of signSGD. The convergence of the proposed method is established in the presence of arbitrary data heterogeneity. The Byzantine resilience of sign-based gradient descent methods is quantified, and the error-feedback mechanism is further incorporated to boost the learning performance Experimental results on the MNIST dataset, the CIFAR-10 dataset, and the Tiny-ImageNet dataset corroborate the effectiveness of the proposed methods.

Talk to us

Join us for a 30 min session where you can share your feedback and ask us any queries you have

Schedule a call

R Discovery Prime

R Discovery Prime

Sign-Based Gradient Descent With Heterogeneous Data: Convergence and Byzantine Resilience.

Abstract

Talk to us

Similar Papers

More From: IEEE transactions on neural networks and learning systems

Lead the way for us

Similar Papers

Direct Estimation of Weights and Efficient Training of Deep Neural Networks without SGD
Nima Dehmamy ... Neda Rohani
-
Nima Dehmamy, et. al.Nima Dehmamy ... Neda Rohani
01 May 2019
01 May 2019

Investigations on hessian-free optimization for cross-entropy training of deep neural networks
Simon Wiesler ... Jinyu Li
-
Simon Wiesler, et. al.Simon Wiesler ... Jinyu Li
25 Aug 2013
25 Aug 2013

Byzantine Resilience With Reputation Scores
Jayanth Regatti ... Abhishek Gupta
-
Jayanth Regatti, et. al.Jayanth Regatti ... Abhishek Gupta
27 Sep 2022
27 Sep 2022

Genetic CFL: Hyperparameter Optimization in Clustered Federated Learning.
Shaashwat Agrawal ... Praveen Kumar Reddy Maddikunta
Computational Intelligence and Neuroscience | VOL. 2021
Shaashwat Agrawal, et. al.Shaashwat Agrawal ... Praveen Kumar Reddy Maddikunta
01 Jan 2020
Computational Intelligence and Neuroscience | VOL. 2021

Editage

Paperpal

R Discovery

Mind the Graph

R Discovery Prime

R Discovery Prime

Sign-Based Gradient Descent With Heterogeneous Data: Convergence and Byzantine Resilience.

Abstract

Talk to us

Similar Papers

More From: IEEE transactions on neural networks and learning systems