西安交通大学计算机科学与技术系,西安,710049
网络首发:2018-12-10,
纸质出版:2018
移动端阅览
刘闯, 何峰, 肖兮, 等. 计算流体力学程序单核指令级优化方法[J]. 西安交通大学学报, 2018,52(12):77-83.
LIU Chuang, HE Feng, XIAO Xi, et al. A Single-Core Instruction-Level Optimization Method for Computational Fluid Dynamics Programs[J]. 2018, 52(12): 77-83.
刘闯, 何峰, 肖兮, 等. 计算流体力学程序单核指令级优化方法[J]. 西安交通大学学报, 2018,52(12):77-83. DOI: 10.7652/xjtuxb201812012.
LIU Chuang, HE Feng, XIAO Xi, et al. A Single-Core Instruction-Level Optimization Method for Computational Fluid Dynamics Programs[J]. 2018, 52(12): 77-83. DOI: 10.7652/xjtuxb201812012.
针对目前大多数计算流体力学程序对系统的单核计算能力利用不足
提出一种针对计算流体力学程序的单核指令级优化方法。该方法首先分析程序的性能指标存在潜在的性能不足
根据分析结果进行优化; 依据容器的存储特性和系统的访存特性
对程序的存储结构和访存顺序进行调整
以优化空间开销和访存性能; 对CPU的流水机制进行分析
在循环和分支中消除指令的控制相关和数据相关从而达到减少流水中断率的目的; 分析编译器对高级语言的处理特点并结合系统中的运行时栈在指令级作出分析
优化指令结构从而减少指令冗余和降低指令复杂度。实验结果表明
在TIANHE-1A超级计算机系统上进行测试
与优化前程序相比
优化后的程序执行时间约减少68.34%
空间消耗约减少55.43%。通过对程序性能各项指标进行分析的结果表明
程序在流水中断率、缓存命中率及机器指令数等性能指标上均有大幅地提升
该方法优化覆盖范围多于目前其他优化方法
有较好的优化效果
在计算流体力学程序优化研究中具有一定的借鉴价值。
A single-core instruction-level optimization method for computational fluid dynamics(CFD)programs is proposed to overcome the shortage of most current CFD programs in utilizing the single-core computing power of a system. The method first analyzes the performance indicators of a program to find potential performance deficiencies
and then optimizes them according to the analysis results. Then the memory structure and memory access sequence are adjusted according to the memory access characteristics of the system and the container to optimize memory access performance and space overhead. The pipeline mechanism of CPU is analyzed
and the control correlation and data correlation of instructions in loops and branches are eliminated to reduce the pipeline interruption rate. Both the characteristics of a compiler processing the high-level language and the runtime stack at the instruction level are analyzed to optimize the instruction structure and to reduce instruction redundancy and duplication. Experimental results show that the performance of the optimized program is greatly improved. Testings on TIANHE-1A supercomputer system show that the execution time of the program reduces by 68.34% and the space consumption reduces by 55.43%. Analyses show that the performance of the program is greatly improved in pipeline interruption rate
cache hit rate and number of machine instructions. It shows that the proposed method has more coverage than other existing optimization methods and better optimization effect
and has a good reference value.
ASTORGA D D R, DOLZ M F, SANCHEZ L M. Discovering pipeline parallel patterns in sequential legacy C++ codes [C]∥Proceedings of the International Workshop on Programming Models and Applications for Multicores and Manycores. New York, USA: ACM, 2016: 11-19
TAYLOR B, MARCO V S, WANG Zheng, et al. Adaptive optimization for OpenCL programs on embedded heterogeneous systems [C]∥Proceedings of the 18th ACM SIGPLAN/SIGBED Conference on Languages, Compilers, and Tools for Embedded Systems. New York, USA: ACM, 2017: 11-20
BIFERALE L, MANTOVANi F, PIVANTI M, et al. Optimization of multi-phase compressible lattice Boltzmann codes on massively parallel multi-core systems [J]. Procedia Computer Science, 2011, 4(4): 994-1003
车永刚, 张理论, 王勇献, 等. 一个结构网格并行CFD程序的单机性能优化 [J]. 计算机科学, 2013, 40(3): 116-120.CHE Yonggang, ZHANG Lilun, WANG Yongxian, et al. Single-machine performance optimization of a structured grid parallel CFD program [J]. Computer Science, 2013, 40(3): 116-120
NGUYEN A T, REITER S, RIGO P. A review on simulation-based optimization methods applied to building performance analysis [J]. Applied Energy, 2014, 113(6): 1043-1058
ALKHANAK E N, LEE S P, REZAEI R, et al. Cost optimization approaches for scientific workflow scheduling in cloud and grid computing: a review, classifications, and open issues [J]. Journal of Systems & Software, 2016, 113(5): 1-26
XU Jie, HUANG E, CHEN C H, et al. Simulation optimization: a review and exploration in the new era of cloud computing and big data [J]. Asia-Pacific Journal of Operational Research, 2015, 32(3): 11-34
CHE Y. Optimization of a parallel CFD code and its performance evaluation on Tianhe-1A [J]. Computing & Informatics, 2015, 33(6): 1377-1399
罗红兵, 张晓霞, 王伟, 等. 科学计算应用程序单核指令级优化研究 [J]. 计算机研究与发展, 2014, 51(6): 1263-1269.LUO Hongbing, ZHANG Xiaoxia, WANG Wei, et al. Instruction level parallel optimizing for scientific computing application [J]. Journal of Computer Research & Development, 2014, 51(6): 1263-1269
张宝印, 莫则尧, 曹小林. 基于计算缓存方法的分子动力学程序性能优化 [J]. 计算机工程与科学, 2009, 31(11): 77-79.ZHANG Baoyin, MO Zeyao, CAO Xiaolin. Molecular dynamics program performance optimization based on computational buffer method [J]. Computer Engineering and Science, 2009, 31(11): 77-79
SIDDIQUE N A, GRUBEL P A, BADAWY A H A, et al. A performance study of the time-varying cache behavior: a study on APEX, Mantevo, NAS, and PARSEC [J]. Journal of Supercomputing, 2018, 74(2): 665-695
DING W, KANDEMIR M. CApRI: Cache-conscious data reordering for irregular codes [J]. ACM Sigmetrics Performance Evaluation Review, 2014, 42(1): 477-489
贺爱香, 顾乃杰, 苏俊杰. 基于多核ARM体系结构的基础函数优化方法 [J]. 计算机工程, 2018, 44(5): 47-52.HE Aixiang, GU Naijie, SU Junjie. Fundamental function optimization method based on multi-core ARM architecture [J]. Computer Engineering, 2018, 44(5): 47-52
OYARZUN G, BORRELL R, GOROBETS A, et al. Efficient CFD code implementation for the ARM-based Mont-Blanc architecture [J]. Future Generation Computer Systems, 2017, 79(3): 786-796
申小伟, 叶笑春, 王达, 等. 一种面向科学计算的数据流优化方法 [J]. 计算机学报, 2017, 40(9): 2181-2196.SHEN Xiaowei, YE Xiaochun, WANG Da, et al. A data flow optimization method for scientific computing [J]. Journal of Computer Science, 2017, 40(9): 2181-2196
ZHANG Chuhua, MIAO Yongmiao, GU Chuangang. Numerical simulations of three dimensional turbulent flows in a shrouded backswept impeller at design and off-design flow rates using unstructured grid method [C]∥Proceedings of the ASME Turbo Expo: Power for Land, Sea, and Air. New York: ASME, 2000: V001T03A033
ZHAO Lei, ZHANG Chuhua. A parallel unstructured finite volume method for all speed flows [J]. Numerical Heat Transfer: Part B Fundamentals, 2014, 65(4): 336-358.
0
浏览量
6
下载量
0
CSCD
关联资源
相关文章
相关作者
相关机构
京公网安备11010802024621