A novel data partitioning method is proposed to address the problem that the "CPU+GPU" heterogeneous parallel processing system cannot fully utilize its resources when average-partition data blocks in batches is processed to deal with the extra overhead for communication. Application data is processed by GPU after being partitioned into blocks with different sizes in proportion by taking the communication bandwidth and the GPU computing capacity into account. Therefore
PCI-E bus and GPU can work in parallel in a period of time to overlap communication and computation. The partitioned data blocks can utilize system resources as much as possible
and hence the mutual waiting time between data transferring and computing can be reduced. Experimental results show that application performance is raised significantly by effectively overlapping communication and computation. Comparisons with no-partition and average-partition show that matrix multiplication's performance is improved by about 5% and 3%
while Fast Fourier Transform's performance is enhanced by about 30% and 6%
WU Enhua. State of the art and future challenge on general purpose computation by graphics processing unit[J]. Journal of Software,2004,15(10):1493-1504.
NVIDIA Corporation. NVIDIA CUDA programming guide [EB/OL] [2010-07-15]. http:∥www.nvidia.com/object/cuda_home_new.html.
FENG Guofu, DONG Xiaoshe, DING Yanfei, et al. A memory access technology of heterogeneous multi-core system based on cell broadband engine architecture[J]. Journal of Xi'an Jiaotong University, 2009, 43(2):1-5.
ZHOU Guoliang, CHEN Hong, LI Cuiping, et al. Parallel data cube computation on graphic processing units[J]. Chinese Journal of Computers, 2010,33(10):1788-1808.
YANG Yang, RAART K V, CASANOVA H. Multiround algorithms for scheduling divisible loads[J]. IEEE Transactions on Parallel and Distributed Systems, 2005,16(11):1092-1102.
TAO Yongcai, JIN Hai, WU Song, et al. Adaptive multi-round scheduling strategy for divisible workloads in grid environments[C]∥Proceedings of the 23rd International Conference on Information Networking. New York, USA:ACM, 2009:260-264.
SHET G A, SADAYAPPAN P, BERNHOLDT E D, et al. A framework for characterizing overlap of communication and computation in parallel applications[J]. Cluster Computing, 2008,11(1): 75-90.
ANTHONY D, LORI P, MARTIN S. MPI-aware compiler optimizations for improving communication-computation overlap[C]∥Proceedings of the 23th International Conference on Supercomputing. New York,USA: ACM, 2009: 316-325.
ZHOU Yongbin,ZHANG Junchao,ZHANG Shuai,et al. Software/hardware co-design for 1-D FFT optimization on many-core architecture[J]. Chinese Journal of Computers, 2008,31(11):2005-2014.