

浏览全部资源
扫码关注微信
兰州交通大学电子与信息工程学院,兰州,730070
Published:2024
移动端阅览
JU Tao, KANG Heting, LIU Shuai, et al. A Dynamic Layer-Wise Gradient Sparsity and Gradient Merging Optimization Method for Deep Neural Networks[J]. 2024, 58(9): 105-116.
JU Tao, KANG Heting, LIU Shuai, et al. A Dynamic Layer-Wise Gradient Sparsity and Gradient Merging Optimization Method for Deep Neural Networks[J]. 2024, 58(9): 105-116. DOI: 10.7652/xjtuxb202409011.
针对数据并行方法加速大规模深度神经网络时易出现的通信开销大、训练耗时长、资源利用率不高的问题
提出了一种深度神经网络动态分层梯度稀疏化及梯度合并优化方法。首先
将梯度稀疏化压缩与流水线并行技术相结合
提出动态分层梯度稀疏优化方法
为每层神经网络匹配一个合适的阈值
通过在后续迭代时动态调整该阈值
实现对每层网络传输梯度的自适应压缩。然后
提出了层梯度合并方法
利用动态规划算法对层梯度合并时的通信开销、稀疏化及层梯度计算时间进行权衡优化
求解出最佳的层梯度合并组合
并将多层小尺度梯度张量合并为一层通信
以降低分层梯度决策时引入的过高通信延迟开销。最后
将求解出的最佳层梯度合并组合应用于具体的训练迭代过程。实验结果表明:与已有方法相比
所提方法可在保证模型训练精度的同时大大降低通信开销
提升模型的训练速度; 与未压缩方法相比
训练速度最大可提升1.99倍。
A dynamic layer-wise gradient sparsity and gradient aggregation optimization strategy for deep neural networks is proposed to address the challenges posed by substantial communication overhead
prolonged training duration
and suboptimal resource utilization associated with the acceleration of large-scale deep neural networks through data parallelism. Initially
a dynamic layer-wise gradient sparsity optimization method is proposed by combining gradient sparsity compression with pipeline parallelism. Each neural network layer is assigned an appropriate threshold
which is adjusted dynamically in subsequent iterations to achieve adaptive compression of gradient transmission for each layer. Subsequently
a layer-wise gradient merging method is introduced. Leveraging dynamic programming
this method optimizes communication overhead
sparsity
and layer gradient computation time during layer-wise gradient merging
determining the optimal combination for merging multiple layers of small-scale gradient tensors into a single communication layer. This aims to reduce the high communication latency introduced during layer-wise gradient decision-making. Finally
the determined optimal layer-wise gradient merging combination is applied to the specific training iteration process. Experimental results demonstrate that the proposed method
compared to existing methods
significantly reduces communication overhead and enhances model training speed while ensuring model training accuracy. It achieves a maximum training speed up of 1.99 times compared to the uncompressed method.
朱泓睿, 元国军, 姚成吉, 等. 分布式深度学习训练网络综述 [J]. 计算机研究与发展, 2021, 58(1): 98-115.
ZHU Hongrui, YUAN Guojun, YAO Chengji, et al. Survey on network of distributed deep learning training [J]. Journal of Computer Research and Development, 2021, 58(1): 98-115.
王帅, 李丹. 分布式机器学习系统网络性能优化研究进展 [J]. 计算机学报, 2022, 45(7): 1384-1411.
WANG Shuai, LI Dan. Research progress on network performance optimization of distributed machine learning system [J]. Chinese Journal of Computers, 2022, 45(7): 1384-1411.
巨涛, 赵宇阳, 刘帅, 等. 面向图片识别的深度学习模型并行优化方法 [J]. 西安交通大学学报, 2023, 57(1): 141-151.
JU Tao, ZHAO Yuyang, LIU Shuai, et al. A parallel optimization method of deep learning model for image recognition [J]. Journal of Xi'an Jiaotong University, 2023, 57(1): 141-151.
LI Mu, ANDERSEN D G, PARK J W, et al. Scaling distributed machine learning with the parameter server [C]//Proceedings of the 11th USENIX Conference on Operating Systems Design and Implementation. Berkeley, CA, USA: USENIX Association, 2014: 583-598.
朱虎明, 李佩, 焦李成, 等. 深度神经网络并行化研究综述 [J]. 计算机学报, 2018, 41(8): 1861-1881.
ZHU Huming, LI Pei, JIAO Licheng, et al. Review of parallel deep neural network [J]. Chinese Journal of Computers, 2018, 41(8): 1861-1881.
BOTTOU L, CURTIS F E, NOCEDAL J, et al. Optimization methods for large-scale machine learning [J]. SIAM Review, 2018, 60(2): 223-311.
高赫然, 吴恒, 许源佳, 等. 面向深度学习训练的内存交换机制综述 [J]. 软件学报, 2023, 34(12): 5862-5886.
GAO Heran, WU Heng, XU Yuanjia, et al. Survey on memory swapping mechanism for deep learning training [J]. Journal of Software, 2023, 34(12): 5862-5886.
巨涛, 刘帅, 王志强, 等. 深度神经网络模型任务切分及并行优化方法 [J/OL]. 北京航空航天大学学报: 1-18[2024-04-13]. https://doi.org/10.13700/j.bh.1001-5965.2022.0731.
JU Tao, LIU Shuai, WANG Zhiqiang, et al. Task segmentation and parallel optimization of DNN model [J/OL]. Journal of Beijing University of Aeronautics and Astronautics: 1-18[2024-04-13]. https://doi.org/10.13700/j.bh.1001-5965.2022.0731.
YAN Guangfeng, LI Tan, WU Kui, et al. Killing two birds with one stone: quantization achieves privacy in distributed learning [J]. Digital Signal Processing, 2024, 146: 104353.
陈世达, 刘强, 韩亮. 降低分布式训练通信的梯度稀疏压缩方法 [J]. 浙江大学学报(工学版), 2021, 55(2): 386-394.
CHEN Shida, LIU Qiang, HAN Liang. Gradient sparsification compression approach to reducing communication in distributed training [J]. Journal of Zhejiang University(Engineering Science), 2021, 55(2): 386-394.
王恩东, 闫瑞栋, 郭振华, 等. 分布式训练系统及其优化算法综述 [J]. 计算机学报, 2024, 47(1): 1-28.
WANG Endong, YAN Ruidong, GUO Zhenhua, et al. A survey of distributed training system and its optimization algorithms [J]. Chinese Journal of Computers, 2024, 47(1): 1-28.
STROM N. Scalable distributed DNN training using commodity GPU cloud computing [C]//Proceedings of the Interspeech 2015. Paris, France: International Speech Communication Association, 2015: 1488-1492.
DRYDEN N, MOON T, JACOBS S A, et al. Communication quantization for data-parallel training of deep neural networks [C]//2016 2nd Workshop on Machine Learning in HPC Environments(MLHPC). Piscataway, NJ, USA: IEEE, 2016: 1-8.
CHEN C Y, CHOI J, BRAND D, et al. AdaComp: adaptive residual gradient compression for data-parallel distributed training [C]//Proceedings of the Thirty-Second AAAI Conference on Artificial Intelligence and Thirtieth Innovative Applications of Artificial Intelligence Conference and Eighth AAAI Symposium on Educational Advances in Artificial Intelligence. Palo Alto, CA, USA: AAAI Press, 2018: 2827-2835.
RENGGLI C, ASHKBOOS S, AGHAGOLZADEH M, et al. SparCML: high-performance sparse communication for machine learning [C]//Proceedings of the International Conference for High Performance Computing, Networking, Storage and Analysis. New York, USA: Association for Computing Machinery, 2019: 11.
SATTLER F, WIEDEMANN S, MÜLLER K R, et al. Sparse binary compression: towards distributed deep learning with minimal communication [C]//2019 International Joint Conference on Neural Networks(IJCNN). Piscataway, NJ, USA: IEEE, 2019: 1-8.
SHI Shaohuai, WANG Qiang, ZHAO Kaiyong, et al. A distributed synchronous SGD algorithm with global top-k sparsification for low bandwidth networks [C]//2019 IEEE 39th International Conference on Distributed Computing Systems(ICDCS). Piscataway, NJ, USA: IEEE, 2019: 2238-2247.
LIN Yujun, HAN Song, MAO Huizi, et al. Deep gradient compression: reducing the communication bandwidth for distributed training [EB/OL].(2020-06-23)[2024-01-01]. https://arxiv.org/abs/1712.01887.
SHI Shaohuai, TANG Zhenheng, WANG Qiang, et al. Layer-wise adaptive gradient sparsification for distributed deep learning with convergence guarantees [M]//Frontiers in Artificial Intelligence and Applications. Amsterdam, Netherlands: IOS Press, 2020: 1467-1474.
SHI Shaohuai, CHU Xiaowen, CHEUNG K C, et al. Understanding top-K sparsification in distributed deep learning [EB/OL].(2019-11-20)[2024-01-01]. https://arxiv.org/abs/1911.08772.
GUO Bingjun, LIU Yazhi, ZHANG Chunyang. A partition based gradient compression algorithm for distributed training in AIoT [J]. Sensors, 2021, 21(6): 1943.
ABDELMONIEM A M, AHMED A, ALOUINI M S, et al. An efficient statistical-based gradient compression technique for distributed training systems [J]. Proceedings of Machine Learning and Systems, 2021, 3: 297-322.
CHMIEL B, BEN-URI L, SHKOLNIK M, et al. Neural gradients are near-lognormal: improved quantized and sparse training [EB/OL].(2020-10-12)[2024-01-01]. https://arxiv.org/abs/2006.08173.
SONG Yingjie, AI Yongbao, XIAO Xiong, et al. HCEC: an efficient geo-distributed deep learning training strategy based on wait-free back-propagation [J]. Journal of Systems Architecture, 2024, 148: 103070.
ZHANG Hao, ZHENG Zeyu, XU Shizhen, et al. Poseidon: an efficient communication architecture for distributed deep learning on GPU clusters [C]//Proceedings of the 2017 USENIX Conference on Usenix Annual Technical Conference. Berkeley, CA, USA: USENIX Association, 2017: 181-193.
JAYARAJAN A, WEI Jinliang, GIBSON G, et al. Priority-based parameter propagation for distributed DNN training [J]. Proceedings of Machine Learning and Systems, 2019, 1: 132-145.
PENG Yanghua, ZHU Yibo, CHEN Yangrui, et al. A generic communication scheduler for distributed DNN training acceleration [C]//Proceedings of the 27th ACM Symposium on Operating Systems Principles. New York, USA: Association for Computing Machinery, 2019: 16-29.
吴艳霞, 梁楷, 刘颖, 等. 深度学习FPGA加速器的进展与趋势 [J]. 计算机学报, 2019, 42(11): 2461-2480.
WU Yanxia, LIANG Kai, LIU Ying, et al. The progress and trends of FPGA-based accelerators in deep learning [J]. Chinese Journal of Computers, 2019, 42(11): 2461-2480.
0
Views
5
下载量
0
CSCD
Publicity Resources
Related Articles
Related Author
Related Institution
京公网安备11010802024621