Towards accelerating model parallelism in distributed deep learning systems.

Hyeonseong Choi,Byung Hyun Lee,Jaehwan Lee,Se Young Chun

doi:10.1371/journal.pone.0293338

Hyeonseong Choi, Byung Hyun Lee + Show 2 more

Open Access

https://doi.org/10.1371/journal.pone.0293338

Copy DOI

Abstract

Modern deep neural networks cannot be often trained on a single GPU due to large model size and large data size. Model parallelism splits a model for multiple GPUs, but making it scalable and seamless is challenging due to different information sharing among GPUs with communication overhead. Specifically, we identify two key issues to make the parallelism being inefficient and inaccurate; an efficient pipelining technique is crucial to maximize GPU utilization and normalizations in deep neural networks may affect the performance due to different statistics sharing of mini-batch. In this work, we address these issues by investigating efficient pipelining for model parallelism and effective normalizations in model / data parallelisms when training a model with large mini-batch in multiple GPUs so that the model performance in accuracy can not be compromised. Firstly, we propose a novel method to search for an optimal micro-batch size considering the number of GPUs and memory size for model parallelism. For efficient pipelining, mini-batch is usually divided into smaller batches (called micro-batch). To maximize the utilization of GPU computing resources, training should be performed with the optimal micro-batch size. Our proposed micro-batch size search algorithm achieved increased image throughput by up to 12% and improved trainable mini-batch size by 25% as compared to the conventional model parallelism method. Secondly, we investigate normalizations in distributed deep learning training for different parallelisms. Our experiments using different normalization methods suggested that the performance with batch normalization can be improved by sharing the batch information among GPUs when performing data parallelism. It was also confirmed that group normalization helped minimizing accuracy degradation when performing model parallelism with pipelining and yielded consistent accuracies for diverse mini-batch sizes.

Talk to us

Join us for a 30 min session where you can share your feedback and ask us any queries you have

Schedule a call

R Discovery Prime

R Discovery Prime

Towards accelerating model parallelism in distributed deep learning systems.

Abstract

Talk to us

Similar Papers

More From: PloS one

Lead the way for us

Journal: PloS one	Publication Date: Nov 2, 2023
License type: CC BY 4.0

Similar Papers

Optimizing makespan and resource utilization for multi-DNN training in GPU cluster
Zhongjin Li ... Francesco Piccialli
Future Generation Computer Systems | VOL. 125
Zhongjin Li, et. al.Zhongjin Li ... Francesco Piccialli
24 Jun 2021
Future Generation Computer Systems | VOL. 125

ScaleDNN: Data Movement Aware DNN Training on Multi-GPU
Weizheng Xu ... Xulong Tang
-
Weizheng Xu, et. al.Weizheng Xu ... Xulong Tang
01 Nov 2021
01 Nov 2021

A Comprehensive Evaluation of Well Completion and Production Performance in Bakken Shale Using Data-Driven Approaches
Shuhua Wang ... Shengnan Chen
-
Shuhua Wang, et. al.Shuhua Wang ... Shengnan Chen
24 Aug 2016
24 Aug 2016

Filter Response Normalization Layer: Eliminating Batch Dependence in the Training of Deep Neural Networks
Saurabh Singh ... Shankar Krishnan
-
Saurabh Singh, et. al.Saurabh Singh ... Shankar Krishnan
01 Jun 2020
01 Jun 2020

Editage

Paperpal

R Discovery

Mind the Graph

R Discovery Prime

R Discovery Prime

Towards accelerating model parallelism in distributed deep learning systems.

Abstract

Talk to us

Similar Papers

More From: PloS one