A profiling based optimization method for CPU-GPU heterogeneous parallel processing system is proposed to address the problem that the present optimization strategies get sectional thus failed to guide a global optimization. It is composed of the optimization strategy library
the profiling tool library
and the strategy deploy module
and the optimization strategy library divides the performance promotion process into a three-level optimization
including the memory-access level
the kernel-speedup level
and the data-partition level. The profiling tool library realizes three-level profiling mechanisms towards three-level optimizations to obtain application information
and the strategy deploy module guides users to choose an adaptive strategy with the information obtained by profiling tool library. Experimental results show that the proposed one is able to guide the optimization process of applications transplanted to heterogeneous parallel system. The performance for matrix multiplication and fast Fourier transform are improved obviously
and the final performance is heightened by 30% compared with the memory-level optimization.
WU Enhua. State of the art and future challenge on general purpose computation by graphics processing unit[J]. Journal of Software, 2004, 15(10): 1493-1504.
RYOO S, RODRIGUES C I, BAGHSORKHI S S, et al. Optimization principles and application performance evaluation of a multithreaded GPU using CUDA[C]∥Proceedings of the 13th ACM SIGPLAN Symposium on Principles and Practice of Parallel Programming. New York, USA: ACM, 2008:73-82.
RYOO S, RODRIGUES C I, STONE S S, et al. Program optimization space pruning for a multithreaded GPU[C]∥Proceedings of the Sixth Annual IEEE/ACM International Symposium on Code Generation and Optimization. Boston, USA: ACM, 2008: 195-204.
ZHANG EDDY Z, JIANG Yunlian, GUO Ziyu, et al. Streamlining GPU applications on the fly: thread divergence elimination through runtime thread-data remapping[C]∥Proceedings of the 24th ACM International Conference on Supercomputing. New York, USA: ACM, 2010: 115-126.
ZHANG Bao, CAO Haijun, DONG Xiaoshe, et al. Novel GPU data partitioning method to overlap communication and computation[J]. Journal of Xi'an Jiaotong University, 2011, 45(4): 1-6.
YANG Yi, XIANG Ping, KONG Jingfei, et al. A GPGPU compiler for memory optimization and parallelism management[C]∥Proceedings of the 2010 ACM SIGPLAN Conference on Programming Language Design and Implementation. New York, USA: ACM, 2010: 86-97.
MALONY A D, BIERSDORFF S, MAYANGLAMBAM S. An experimental approach to performance measurement of heterogeneous parallel applications using CUDA[C]∥Proceedings of the 24th ACM International Conference on Supercomputing. New York, USA: ACM, 2010: 127-136.
BAGHSORKHI S S, DELAHAYE M, PATEL S J, et al. An adaptive performance modeling tool for GPU architectures[C]∥Proceedings of the 15th ACM SIGPLAN Symposium on Principles and Practice of Parallel Programming. New York, USA: ACM, 2010: 105-114.
NVIDIA Corporation. NVIDIA CUDA Programming guide[EB/OL]. [2010-07-15].http:∥www.nvidia.com/object/cuda_home_new.html.
DONG Xiaoshe, FENG Guofu, WANG Xuhao, et al. Research on multilayer runtime library technology for Cell/B.E. processor[J]. Journal of Computer Research and Development, 2010, 47(4): 561-570.