西安交通大学电子与信息工程学院,西安,710049
网络首发:2012-02-10,
纸质出版:2012
移动端阅览
张保, 董小社, 白秀秀, 等. CPU-GPU系统中基于剖分的全局性能优化方法[J]. 西安交通大学学报, 2012,46(2):17-23.
Profiling Based Optimization Method for CPU-GPU Heterogeneous Parallel Processing System[J]. 2012, 46(2): 17-23.
针对将应用移植到CPU-GPU异构并行系统上时优化策略各自分散、没有一个全局的指导思想的问题
提出了一种基于剖分的全局性能优化方法.该方法由优化策略库、剖分工具库和策略配置模块组成.优化策略库将应用移植到异构并行系统上的性能优化过程划分为访存级、内核加速级和数据划分级3级优化; 针对3级优化剖分工具库提供了3级剖分机制
通过运行时的剖分技术获取剖分信息; 策略配置模块根据所获取的信息指导用户在每级优化中选择合适的优化策略.实验证明
基于剖分的全局性能优化方法可以明确地指导将应用移植到CPU-GPU异构并行系统上的全局优化过程
利用该优化方法后
以矩阵相乘和傅里叶变换为例的应用性能提升明显
最终性能相对于访存级优化最高可提高30%左右.
A profiling based optimization method for CPU-GPU heterogeneous parallel processing system is proposed to address the problem that the present optimization strategies get sectional thus failed to guide a global optimization. It is composed of the optimization strategy library
the profiling tool library
and the strategy deploy module
and the optimization strategy library divides the performance promotion process into a three-level optimization
including the memory-access level
the kernel-speedup level
and the data-partition level. The profiling tool library realizes three-level profiling mechanisms towards three-level optimizations to obtain application information
and the strategy deploy module guides users to choose an adaptive strategy with the information obtained by profiling tool library. Experimental results show that the proposed one is able to guide the optimization process of applications transplanted to heterogeneous parallel system. The performance for matrix multiplication and fast Fourier transform are improved obviously
and the final performance is heightened by 30% compared with the memory-level optimization.
吴恩华. 图形处理器用于通用计算的技术、现状及其挑战[J].软件学报,2004,15(10):1493-1504.
WU Enhua. State of the art and future challenge on general purpose computation by graphics processing unit[J]. Journal of Software, 2004, 15(10): 1493-1504.
RYOO S, RODRIGUES C I, BAGHSORKHI S S, et al. Optimization principles and application performance evaluation of a multithreaded GPU using CUDA[C]∥Proceedings of the 13th ACM SIGPLAN Symposium on Principles and Practice of Parallel Programming. New York, USA: ACM, 2008:73-82.
RYOO S, RODRIGUES C I, STONE S S, et al. Program optimization space pruning for a multithreaded GPU[C]∥Proceedings of the Sixth Annual IEEE/ACM International Symposium on Code Generation and Optimization. Boston, USA: ACM, 2008: 195-204.
ZHANG EDDY Z, JIANG Yunlian, GUO Ziyu, et al. Streamlining GPU applications on the fly: thread divergence elimination through runtime thread-data remapping[C]∥Proceedings of the 24th ACM International Conference on Supercomputing. New York, USA: ACM, 2010: 115-126.
张保,曹海军,董小社,等. 面向图形处理器重叠通信与计算的数据划分方法[J]. 西安交通大学学报,2011,45(4):1-6.
ZHANG Bao, CAO Haijun, DONG Xiaoshe, et al. Novel GPU data partitioning method to overlap communication and computation[J]. Journal of Xi'an Jiaotong University, 2011, 45(4): 1-6.
YANG Yi, XIANG Ping, KONG Jingfei, et al. A GPGPU compiler for memory optimization and parallelism management[C]∥Proceedings of the 2010 ACM SIGPLAN Conference on Programming Language Design and Implementation. New York, USA: ACM, 2010: 86-97.
MALONY A D, BIERSDORFF S, MAYANGLAMBAM S. An experimental approach to performance measurement of heterogeneous parallel applications using CUDA[C]∥Proceedings of the 24th ACM International Conference on Supercomputing. New York, USA: ACM, 2010: 127-136.
BAGHSORKHI S S, DELAHAYE M, PATEL S J, et al. An adaptive performance modeling tool for GPU architectures[C]∥Proceedings of the 15th ACM SIGPLAN Symposium on Principles and Practice of Parallel Programming. New York, USA: ACM, 2010: 105-114.
NVIDIA Corporation. NVIDIA CUDA Programming guide[EB/OL]. [2010-07-15].http:∥www.nvidia.com/object/cuda_home_new.html.
董小社,冯国富,王旭昊,等. 基于Cell多核处理器的层次化运行时支持技术[J]. 计算机研究与发展,2010,47(4):561-570.
DONG Xiaoshe, FENG Guofu, WANG Xuhao, et al. Research on multilayer runtime library technology for Cell/B.E. processor[J]. Journal of Computer Research and Development, 2010, 47(4): 561-570.
应用回归分析的数据关联算法. 西安交通大学学报, 2011,45(8):92-96.
面向图形处理器重叠通信与计算的数据划分方法. 西安交通大学学报, 2011,45(4):1-4.
海量信息异常检测问题的异常概率排序算法. 西安交通大学学报, 2011,45(4):36-40.
一种面向多处理器系统的在线低功耗调度算法. 西安交通大学学报, 2010,44(8):15-19.
面向服务级别协议的核实与规划服务质量系统框架研究. 西安交通大学学报, 2010,44(6): 1-5.
利用投影时序逻辑的多内核进程调度建模与验证. 西安交通大学学报, 2010,44(3):52-57.
面向属性约束的自动信任协商模型. 西安交通大学学报, 2009,43(8): 1-5.
面向Cell宽带引擎架构的异构多核访存技术. 西安交通大学学报, 2009,43(2):1-5.
基于思维进化的集群作业调度方法研究. 西安交通大学学报, 2008,42(6):651-654.
支持多种Linux版本的动态内核性能测试技术. 西安交通大学学报, 2008,42(6):674-678.
自适应大规模服务器集群监控系统的构建. 西安交通大学学报, 2008,42(4):399-403.
网格测试引擎:一种构建网格测试环境的基楚构架. 西安交通大学学报, 2007,41(8):884-888.
按序选取的并行存储系统数据分布方式研究. 西安交通大学学报, 2007,41(8):899-902.
0
浏览量
4
下载量
4
CSCD
关联资源
相关文章
相关作者
相关机构
京公网安备11010802024621