西安交通大学计算机科学与技术系,西安,710049
纸质出版:2011
移动端阅览
张保 1. 面向图形处理器重叠通信与计算的数据划分方法[J]. 西安交通大学学报, 2011,45(4):1-5+11.
张保 1. Novel GPU Data Partitioning Method to Overlap Communication and Computation[J]. 2011, 45(4): 1-5+11.
针对“主核心+协处理器”式异构并行系统采用数据平均划分再分批执行的方法来解决主协式处理架构的额外通信开销时未能充分利用系统资源的问题
提出了一种新的数据比例划分方法.结合系统通信带宽和图形处理器(GPU)的计算能力
将应用数据按比例划分为大小不同的数据块后分批提交给GPU处理
使系统的传输资源PCI-E总线和计算资源GPU在一段时间内并行工作
从而实现了应用通信与计算的重叠.在处理按照比例划分的数据块过程中
尽可能充分利用系统的传输资源和计算资源
以减少数据传输和计算的相互等待时间.实验结果表明
采用数据比例划分方法后的应用性能明显提高
可以有效地重叠通信与计算时间
矩阵相乘和快速傅里叶变换总执行时间比未划分时分别减少了5%和30%左右
比平均划分时分别减少了3%和6%左右.
A novel data partitioning method is proposed to address the problem that the "CPU+GPU" heterogeneous parallel processing system cannot fully utilize its resources when average-partition data blocks in batches is processed to deal with the extra overhead for communication. Application data is processed by GPU after being partitioned into blocks with different sizes in proportion by taking the communication bandwidth and the GPU computing capacity into account. Therefore
PCI-E bus and GPU can work in parallel in a period of time to overlap communication and computation. The partitioned data blocks can utilize system resources as much as possible
and hence the mutual waiting time between data transferring and computing can be reduced. Experimental results show that application performance is raised significantly by effectively overlapping communication and computation. Comparisons with no-partition and average-partition show that matrix multiplication's performance is improved by about 5% and 3%
while Fast Fourier Transform's performance is enhanced by about 30% and 6%
respectively.
吴恩华. 图形处理器用于通用计算的技术、现状及其挑战[J].软件学报,2004,15(10):1493-1504.
WU Enhua. State of the art and future challenge on general purpose computation by graphics processing unit[J]. Journal of Software,2004,15(10):1493-1504.
冯国富,董小社,丁彦飞,等. 面向Cell宽带引擎架构的异构多核访存技术[J].西安交通大学学报,2009,43(2):1-5.
FENG Guofu, DONG Xiaoshe, DING Yanfei, et al. A memory access technology of heterogeneous multi-core system based on cell broadband engine architecture[J]. Journal of Xi'an Jiaotong University, 2009, 43(2):1-5.
周国亮,陈红,李翠平,等. 基于图形处理器的并行方体计算[J].计算机学报,2010,33(10):1788-1808.
ZHOU Guoliang, CHEN Hong, LI Cuiping, et al. Parallel data cube computation on graphic processing units[J]. Chinese Journal of Computers, 2010,33(10):1788-1808.
YANG Yang, RAART K V, CASANOVA H. Multiround algorithms for scheduling divisible loads[J]. IEEE Transactions on Parallel and Distributed Systems, 2005,16(11):1092-1102.
TAO Yongcai, JIN Hai, WU Song, et al. Adaptive multi-round scheduling strategy for divisible workloads in grid environments[C]∥Proceedings of the 23rd International Conference on Information Networking. New York, USA:ACM, 2009:260-264.
SHET G A, SADAYAPPAN P, BERNHOLDT E D, et al. A framework for characterizing overlap of communication and computation in parallel applications[J]. Cluster Computing, 2008,11(1): 75-90.
ANTHONY D, LORI P, MARTIN S. MPI-aware compiler optimizations for improving communication-computation overlap[C]∥Proceedings of the 23th International Conference on Supercomputing. New York,USA: ACM, 2009: 316-325.
周永彬,张军超,张帅,等. 基于软硬件的协同支持在众核上对1-DFFT算法的优化研究[J]. 计算机学报,2008,31(11):2005-2014.
ZHOU Yongbin,ZHANG Junchao,ZHANG Shuai,et al. Software/hardware co-design for 1-D FFT optimization on many-core architecture[J]. Chinese Journal of Computers, 2008,31(11):2005-2014.
0
浏览量
8
下载量
0
CSCD
关联资源
相关文章
相关作者
相关机构
京公网安备11010802024621