西安交通大学计算机科学与技术学院,710049,西安
西安交通大学网络信息中心,710049,西安
西安交通大学软件学院,710049,西安
作者简介:刘恒(2000—),男,博士生;
王子衡(通信作者),男,助理教授。
收稿:2025-12-10,
纸质出版:2026-07-10
移动端阅览
刘恒, 王子衡, 王强, 等. 采用机器学习的大规模并行程序I/O性能建模与预测方法[J]. 西安交通大学学报, 2026,60(7):207-218.
LIU Heng, WANG Ziheng, WANG Qiang, et al. I/O Performance Modeling and Prediction Method for Large-Scale Parallel Applications Using Machine Learning[J]. Journal of Xi'an Jiaotong University, 2026, 60(7): 207-218.
刘恒, 王子衡, 王强, 等. 采用机器学习的大规模并行程序I/O性能建模与预测方法[J]. 西安交通大学学报, 2026,60(7):207-218. DOI: 10.7652/xjtuxb202607019.
LIU Heng, WANG Ziheng, WANG Qiang, et al. I/O Performance Modeling and Prediction Method for Large-Scale Parallel Applications Using Machine Learning[J]. Journal of Xi'an Jiaotong University, 2026, 60(7): 207-218. DOI: 10.7652/xjtuxb202607019.
针对大规模并行程序因输入/输出(I/O)数据难以获取、建模与优化成本高昂而导致的性能分析效率低下和调优困难问题,提出了一种新的机器学习驱动的大规模并行程序I/O性能建模与预测方法。该方法在建模阶段,一方面基于小规模节点环境下的采样数据,利用线性回归方法构建可用于大规模节点外推的并行程序I/O特征预测模型;另一方面,将小规模节点环境下采集的并行程序I/O特征与对应的I/O栈参数空间共同输入人工神经网络(ANN)进行训练,以学习系统配置参数与I/O性能之间的非线性映射关系,从而构建并行程序的I/O性能预测模型。在预测阶段,使用I/O特征预测模型预测大规模并行程序的I/O特征,并将其与对应的大规模并行程序I/O栈参数空间输入I/O性能预测模型,实现对大规模并行程序I/O性能的准确外推预测。实验结果表明:在国产超级计算机上,采用所提方法在4种典型测试程序IOR、S3D-IO、BT-IO和Flash-IO上的预测精度均较高;当采用1~16节点规模下训练得到的模型对128节点(2048进程)场景进行外推预测时,其平均绝对百分比误差分别为19.07%、18.96%、12.34%和14.16%;基于所构建的I/O性能预测模型对I/O栈参数调优后,4种典型程序I/O性能加速比分别达到16.38、23.16、45.36和65.38倍。
To address the issues of inefficient performance analysis and difficulty in tuning caused by the scarcity of input/output(I/O)data and the prohibitive costs of modeling and optimization in large-scale parallel applications,a novel method for machine learning-driven I/O performance modeling and prediction in large-scale parallel applications is proposed.During the modeling phase,on the one hand,a linear regression-based model for I/O feature prediction in parallel applications is constructed using data sampled from small-scale node environments to enable the extrapolation of features to large-scale configurations;on the other hand,I/O features in parallel applications acquired from small-scale node environments,along with the corresponding I/O stack parameter space,are fed into an artificial neural network(ANN)for training.This allows for the learning of the nonlinear mapping between system configuration parameters and I/O performance,leading to the construction of a model for I/O performance prediction in parallel applications. During the prediction phase,the I/O features of large-scale parallel applications are first predicted via the I/O feature prediction model.These predicted features,together with the corresponding I/O stack parameter space of large-scale parallel applications,are then fed into the I/O performance prediction model to accurately predict the I/O performance of such applications.Experimental results demonstrate that,on a domestic supercomputer,the proposed method exhibits high prediction accuracy across four benchmark applications:IOR,S3D-IO,BT-IO,and Flash-IO.When models trained on 1—16 nodes are extrapolated to a 128-node(2048-process)scenario,the mean absolute percentage errors(MAPEs)are 19.07%,18.96%,12.34%,and 14.16%,respectively.Furthermore,by tuning the I/O stack parameters based on the proposed model,I/O performance speedups of 16.38,23.16,45.36,and 65.38-fold are achieved for the four benchmark applications.
Luu H, Winslett M, Gropp W, et al.A multiplatform study of I/O behavior on petascale supercomputers [C]//Proceedings of the 24th International Symposium on High-Performance Parallel and Distributed Computing.New York, USA:ACM, 2015:33-44.
张成.基于回归分析和集成学习的HPC应用I/O性能优化方法研究[D].西安:西北大学, 2023.
孙经纬.数据驱动的高性能计算程序执行时间预测与优化研究[D].合肥:中国科学技术大学, 2020.
Zhu Zhaobin, Neuwirth S, Lippert T.A comprehensive I/O knowledge cycle for modular and automated HPC workload analysis [C]//2022 IEEE International Conference on Cluster Computing (CLUSTER).Piscataway, NJ, USA:IEEE, 2022:581-588.
Liu Wei, Wu Kai, Liu Jialin, et al.Performance evaluation and modeling of HPC I/O on non-volatile memory [C]//2017 International Conference on Networking, Architecture, and Storage (NAS).Piscataway, NJ, USA:IEEE, 2017:1-10.
Neuwirth S, Paul A K.Parallel I/O evaluation techniques and emerging HPC workloads: a perspective [C]//2021 IEEE International Conference on Cluster Computing (CLUSTER). Piscataway, NJ, USA:IEEE, 2021: 671-679.
汤志航, 兰颢, 刘政国, 等.HiTrain:面向大模型训练的异构内存卸载与I/O优化[J].计算机研究与发展, 2026, 63(3):627-639.
Tang Zhihang, Lan Hao, Liu Zhengguo, et al.Hi-Train:heterogeneous memory offloading and I/O optimization for large language model training [J].Journal of Computer Research and Development, 2026, 63 (3):627-639.
程稳, 李焱, 曾令仿,等.面向Lustre集群存储的应用日志分析及系统自动优化框架[J].计算机工程与科学, 2022, 44(4):594-604.
Cheng Wen, Li Yan, Zeng Lingfang, et al.An application log analysis and system automation optimization framework for Lustre cluster storage [J].Computer Engineering & Science, 2022, 44(4):594-604.
张文韬, 汪璐, 程耀东.基于强化学习的Lustre文件系统的性能调优[J].计算机研究与发展, 2019, 56 (7):1578-1586.
Zhang Wentao, Wang Lu, Cheng Yaodong.Performance optimization of Lustre file system based on reinforcement learning [J].Journal of Computer Research and Development, 2019, 56(7):1578-1586.
田鸿运, 武林平, 董勇, 等.面向大规模集群的并行I/O用户层配置优化策略[J].国防科技大学学报, 2020, 42(2):23-30.
Tian Hongyun, Wu Linping, Dong Yong, et al.Userlevel parallel I/O configuration optimize strategy toward large-scale cluster [J].Journal of National University of Defense Technology, 2020, 42(2):23-30.
Egersdoerfer C, Rashid M H, Dai Dong, et al.Understanding and predicting cross-application I/O interference in HPC storage systems [C]//SC24-W:Workshops of the International Conference for High Performance Computing, Networking, Storage and Analysis. Piscataway, NJ, USA:IEEE, 2024:1330-1339.
Behzad B, Luu H V T, Huchette J, et al.Taming parallel I/O complexity with auto-tuning [C]//Proceedings of the International Conference on High Per formance Computing, Networking, Storage and Analysis.New York, USA:ACM, 2013:68.
Tipu AJ S, ConbhuíPÓ, Howley E.Artificialneural networks based predictions towards the auto-tuning and optimization of parallel IO bandwidth in HPC system [J].Cluster Computing, 2024, 27(1):71-90.
Nicolas L M, Mimouni S, Couvée P, et al.I/O patterns modeling of HPC applications with call stacks for predictive prefetch [J].Future Generation Computer Systems, 2026, 175:108034.
Behzad B, Byna S, Prabhat, et al.Optimizing I/O performance of HPC applications with autotuning [J]. ACM Transactions on Parallel Computing, 2019, 5(4):15.
Bagbaba A, Wang Xuan.Improving the mpi-io performance of applications with genetic algorithm based auto-tuning [C]//2021 IEEE International Parallel and Distributed Processing Symposium Workshops (IPDPSW). Piscataway, NJ, USA:IEEE, 2021:798-805.
Nemirovsky D, Arkose T, Markovic N, et al.A machine learning approach for performance prediction and scheduling on heterogeneous CPUs [C]//201729th International Symposium on Computer Architecture and High Performance Computing (SBAC-PAD).Piscataway, NJ, USA:IEEE, 2017:121-128.
Kim S, Sim A, WU Kesheng, et al.Design and implementation of I/O performance prediction scheme on HPC systems through large-scale log analysis [J]. Journal of Big Data, 2023, 10(1):65.
Liu Zhangyu, Zhang Cheng, Wu Huijun, et al.Optimizing HPC I/O performance with regression analysis and ensemble learning [C]//2023 IEEE International Conference on Cluster Computing (CLUSTER).Piscataway, NJ, USA:IEEE, 2023:234-246.
Meswani M R, Laurenzano M A, Carrington L, et al. Modeling and predicting disk I/O time of HPC applications [C]//2010 DoD High Performance Computing Modernization Program Users Group Conference.Piscataway, NJ, USA:IEEE, 2010:478-486.
Wang Wanxin, Wu Huijun, Yang Lihua, et al.AIO: automating I/O optimization pipeline for data-intensive applications inHPC [C]//2024 IEEE International Symposium on Parallel and Distributed Processing with Applications (ISPA).Piscataway, NJ, USA:IEEE, 2024:1541-1548.
Shan Hongzhang, Shalf J.Using IOR to analyze the I/O performance for HPC platforms [EB/OL].[2015-10-22]. https://escholarship.org/uc/item/9111c60j.
Mendez S, Rexachs D, Luque E.Analyzing the parallel I/O severity of MPI applications [C]//2017 17th IEEE/ACM International Symposium on Cluster, Cloud and Grid Computing (CCGRID).Piscataway, NJ, USA:IEEE, 2017:953-962.
Agarwal M, Singhvi D, Malakar P, et al.Active learning-based automatic tuning and prediction of parallel I/O performance [C]//2019 IEEE/ACM Fourth International Parallel Data Systems Workshop (PD-SW).Piscataway, NJ, USA:IEEE, 2019:20-29.
Rajesh N, Bateman K, Bez J L, et al.TunIO: an AIpowered framework for optimizingHPC I/O [C]//2024 IEEE International Parallel and Distributed Processing Symposium (IPDPS).Piscataway, NJ, USA:IEEE, 2024:494-505.
Chen Si, De Gonzalo S G, Wildani A.Few-shot HPC application runtime prediction [C]//2023 IEEE International Conference on Cluster Computing Workshops (CLUSTER Workshops).Piscataway, NJ, USA:IEEE, 2023: 46-47.
0
浏览量
12
下载量
0
CSCD
关联资源
相关文章
相关作者
相关机构
京公网安备11010802024621