Learning to Learn Gradient Aggregation by Gradient Descent

Jinlong Ji,Lixing Yu,Qianlong Wang,Xuhui Chen,Pan Li

doi:10.24963/ijcai.2019/363

Abstract

In the big data era, distributed machine learning emerges as an important learning paradigm to mine large volumes of data by taking advantage of distributed computing resources. In this work, motivated by learning to learn, we propose a meta-learning approach to coordinate the learning process in the master-slave type of distributed systems. Specifically, we utilize a recurrent neural network (RNN) in the parameter server (the master) to learn to aggregate the gradients from the workers (the slaves). We design a coordinatewise preprocessing and postprocessing method to make the neural network based aggregator more robust. Besides, to address the fault tolerance, especially the Byzantine attack, in distributed machine learning systems, we propose an RNN aggregator with additional loss information (ARNN) to improve the system resilience. We conduct extensive experiments to demonstrate the effectiveness of the RNN aggregator, and also show that it can be easily generalized and achieve remarkable performance when transferred to other distributed systems. Moreover, under majoritarian Byzantine attacks, the ARNN aggregator outperforms the Krum, the state-of-art fault tolerance aggregation method, by 43.14%. In addition, our RNN aggregator enables the server to aggregate gradients from variant local models, which significantly improve the scalability of distributed learning.

Talk to us

Join us for a 30 min session where you can share your feedback and ask us any queries you have

Schedule a call

R Discovery Prime

R Discovery Prime

Learning to Learn Gradient Aggregation by Gradient Descent

Abstract

Talk to us

Similar Papers

Lead the way for us

Similar Papers

DFS: Joint data formatting and sparsification for efficient communication in Distributed Machine Learning
Cheng Yang ... Hongli Xu
Computer Networks | VOL. 229
Cheng Yang, et. al.Cheng Yang ... Hongli Xu
19 Apr 2023
Computer Networks | VOL. 229

Timed Dataflow: Reducing Communication Overhead for Distributed Machine Learning Systems
Peng Sun ... Ta Nguyen Binh Duong
-
Peng Sun, et. al.Peng Sun ... Ta Nguyen Binh Duong
01 Dec 2016
01 Dec 2016

HiPS
Jinkun Geng ... Dan Li
-
Jinkun Geng, et. al.Jinkun Geng ... Dan Li
01 Jan 2018
01 Jan 2018

Online Job Scheduling in Distributed Machine Learning Clusters
Yixin Bao ... Yanghua Peng
-
Yixin Bao, et. al.Yixin Bao ... Yanghua Peng
01 Apr 2018
01 Apr 2018

Editage

Paperpal

R Discovery

Mind the Graph

R Discovery Prime

R Discovery Prime

Learning to Learn Gradient Aggregation by Gradient Descent

Abstract

Talk to us

Similar Papers