兰州交通大学电子与信息工程学院,730070,兰州
作者简介:巨涛(1980—),男,教授,硕士生导师。
收稿:2025-04-14,
纸质出版:2026-02-10
移动端阅览
巨涛, 丁肖健, 郭东雨, 等. 一种面向多模态模型的分区混合并行优化方法[J]. 西安交通大学学报, 2026,60(2):229-240.
JU Tao, DING Xiaojian, GUO Dongyu, et al. A Partitioned Hybrid Parallel Optimization Method for Multimodal Models[J]. Journal of Xi'an Jiaotong University, 2026, 60(2): 229-240.
巨涛, 丁肖健, 郭东雨, 等. 一种面向多模态模型的分区混合并行优化方法[J]. 西安交通大学学报, 2026,60(2):229-240. DOI: 10.7652/xjtuxb202602022.
JU Tao, DING Xiaojian, GUO Dongyu, et al. A Partitioned Hybrid Parallel Optimization Method for Multimodal Models[J]. Journal of Xi'an Jiaotong University, 2026, 60(2): 229-240. DOI: 10.7652/xjtuxb202602022.
针对多模态模型架构复杂、参数量庞大、算力需求高,导致训练难度大、效率低,以及现有数据并行与模型并行策略难以应对内部异构特性的问题,提出了一种面向多模态模型的分区混合并行优化(MMHP)方法。首先根据多模态模型中不同子模块的异构性,构建模块依赖图,识别关键切分点,实现模块分区,保证子模块间的负载均衡;其次,针对不同模块分区的参数规模和计算需求,融合数据并行与模型并行,设计混合并行优化方法以适应异构计算任务的多样化需求;最后,基于动态规划设计计算任务调度优化算法,动态分配计算资源,实现计算任务与资源合理匹配,进一步优化计算资源的利用率,提升模型的训练效率。实验结果表明,在不影响模型训练精度的情况下,所提出的MMHP方法可充分利用计算资源,提升多模态模型的训练效率,与现有主流并行策略相比,训练速度最大可提升2倍。
To address the challenges of complex architecture,large parameter size,and high computational demands in multimodal models,which lead to difficult training and low efficiency,as well as the limitations of existing data and model parallelism strategies in handling internal heterogeneity,a partitioned hybrid parallel optimization(MMHP)method for multimodal models is proposed.First,based on the heterogeneity of different submodules in the multimodal model,a module dependency graph is constructed to identify key partitioning points,achieving module partitioning and ensuring load balancing among submodules.Second,according to the parameter scale and computational requirements of different module partitions,a hybrid parallel optimization method is designed by integrating data parallelism and model parallelism to accommodate the diverse needs of heterogeneous computing tasks.Finally,a computational task scheduling optimization algorithm is developed based on dynamic programming to dynamically allocate computing resources,achieving a reasonable match between computational tasks and resources,further optimizing the utilization of computing resources,and improving model training efficiency. Experimental results show that,without compromising model training accuracy,the proposed MMHP method can fully utilize computing resources and improve the training efficiency of multimodal models.Compared with existing mainstream parallel strategies,the training speed can be increased by up to 2 times.
WANG Jiaqi, JIANG Hanqi, LIU Yiheng, et al.A comprehensive review of multimodal large language models:performance and challenges across different tasks[EB/OL].(2024-08-02)[2024-12-18].https://arxiv.org/abs/2408.01319 .
LIANG Zijing, XU Yanjie, HONG Yifan, et al.A survey of multimodel large language models[C]//Proceedings of the 3rd International Conference on Computer.New York, NY, USA:ACM, 2024:405-409.
LI Jian, LU Weiheng, FEI Hao, et al.A survey on benchmarks of multimodal large language models[EB/OL].(2024-09-06 )[2024-12-19].https://arxiv.org/abs/2408.08632.
ZHANG Min, LI Juntao.A commentary of GPT-3 in-MIT Technology Review 2021[J].Fundamental Research, 2021 , 1(6):831-833.
TOUVRON H, LAVRIL T, IZACARD G, et al. LLaMA:open and efficient foundation language models[EB/OL].(2023-02-27)[2024-12-19].https://arxiv.org/abs/2302.13971 .
HUANG Jun, ZHANG Zhen, ZHENG Shuai, et al. DISTMM:accelerating distributed multimodal model training[C]//Proceedings of the 21st USENIX Symposium on Networked Systems Design and Implementation.USA:USENIX Association, 2024:1157-1171.
王恩东,闫瑞栋,郭振华,等.分布式训练系统及其优化算法综述[J].计算机学报, 2024, 47(1):1-28.
WANG Endong, YAN Ruidong, GUO Zhenhua, et al.A survey of distributed training system and its optimization algorithms[J].Chinese Journal of Computers, 2024, 47(1):1-28.
卢凯,赖志权,李笙维,等.并行智能训练技术:挑战与发展[J].中国科学(信息科学), 2023, 53 (8):1441-1468.
LU Kai, LAI Zhiquan, LI Shengwei, et al.Parallel intelligent computing:development and challenges[J]. Scientia Sinica (Informationis ), 2023, 53 (8 ):1441-1468.
王帅,李丹.分布式机器学习系统网络性能优化研究进展[J].计算机学报, 2022, 45(7):1384-1411.
WANG Shuai, LI Dan.Research progress on network performance optimization of distributed machine learning system[J].Chinese Journal of Computers, 2022, 45(7):1384-1411.
LI Dongsheng, LI Shengwei, LAI Zhiquan, et al.A memory-efficient hybrid parallel framework for deep neural network training[J]. IEEE Transactions on Parallel and Distributed Systems, 2023, 35 (4 ):577-591.
XU Lang, ANTHONY Q, ZHOU Qinghua, et al. Accelerating large language model training with hybrid GPU-based compression[C]//2024 IEEE 24th International Symposium on Cluster, Cloud and Internet Computing (CCGrid).Piscataway, NJ, USA:IEEE, 2024:196-205.
BIAN Song, LI Dacheng, WANG Hongyi, et al.Does compressing activations help model parallel training?[J]. Proceedings of Machine Learning and Systems, 2024, 6:239-252.
CHEN Y C, LI Linjie, YU Licheng, et al.UNITER:UNiversal image-TExt representation learning[C]//Computer Vision-ECCV 2020.Cham, Switzerland:Springer International Publishing, 2020:104-120.
XUE Zhenliang, HU Hanpeng, CHEN Xing, et al. PipeWeaver:addressing data dynamicity in large multimodal model training with dynamic interleaved pipeline[EB/OL].[2025-04-05].https://arxiv.org/abs/2504.14145 .
HUANG Yanping, CHENG Youlong, BAPNA A, et al.GPipe:efficient training of giant neural networks using pipeline parallelism[C]//Proceedings of the 33rd International Conference on Neural Information Processing Systems.Red Hook, NY, USA:Curran Associates Inc., 2019:103-112.
ALAYRAC J B, DONAHUE J, LUC P, et al.Flamingo:a visual language model for few-shot learning[C]//Advances in Neural Information Processing Systems.Red Hook, NY, USA:Curran Associates Inc., 2022:23716-23736.
SHOEYBI M, PATWARY M, PURI R, et al.Megatron-LM:training multi-billion parameter language models using GPU model parallelism[EB/OL].(2020-03-13)[2024-12-26].https://arxiv.org/abs/1909.08053.
OpenAI.GPT-4 technical report[EB/OL].(2024-03-04)[2024-12-28].https://arxiv.org/abs/2303.08774.
巨涛,赵宇阳,刘帅,等.面向图片识别的深度学习模型并行优化方法[J].西安交通大学学报, 2023, 57 (1):141-151.
JU Tao, ZHAO Yuyang, LIU Shuai, et al.A parallel optimization method of deep learning model for image recognition[J].Journal of Xi'an Jiaotong University, 2023, 57(1):141-151.
巨涛,刘帅,火久元,等.深度神经网络模型并行自适应计算任务调度方法[J].吉林大学学报(工学版), 2024, 54(12):3601-3613.
JU Tao, LIU Shuai, HUO Jiuyuan, et al.Adaptive scheduling of computing tasks for deep neural network model parallelism[J].Journal of Jilin University(Engineering and Technology Edition), 2024, 54 (12 ):3601-3613.
巨涛,刘帅,王志强,等.深度神经网络模型任务切分及并行优化方法[J].北京航空航天大学学报, 2024, 50(9):2739-2752.
JU Tao, LIU Shuai, WANG Zhiqiang, et al.Task segmentation and parallel optimization of DNN model[J]. Journal of Beijing University of Aeronautics and Astronautics, 2024, 50(9):2739-2752.
LIANG Feng, ZHANG Zhen, LU Haifeng, et al.Resource allocation and workload scheduling for largescale distributed deep learning:a survey[EB/OL]. (2024-06-12 )[2024-12-28].https://arxiv.org/abs/2406.08115.
杨紫超,吴恒,吴悦文,等.基于性能建模的深度学习训练任务调度综述[J].软件学报, 2025, 36(4):1570-1589.
YANG Zichao, WU Heng, WU Yuewen, et al.Survey on task scheduling of deep learning training based on performance modeling[J].Journal of Software, 2025, 36(4):1570-1589.
RADFORD A, KIM J W, HALLACY C, et al. Learning transferable visual models from natural language supervision[C]//Proceedings of the 38th International Conference on Machine Learning.Chia Laguna Resort, Sardinia, Italy:PMLR, 2021:8748-8763.
LI Junnan, LI Dongxu, XIONG Caiming, et al.BLIP:bootstrapping language-image pre-training for unified vision-language understanding and generation[C]//Proceedings of the 39th International Conference on Machine Learning.Chia Laguna Resort, Sardinia, Italy:PMLR, 2022:12888-12900.
0
浏览量
42
下载量
0
CSCD
关联资源
相关文章
相关作者
相关机构
京公网安备11010802024621