西安交通大学电子与信息工程学院,西安,710049
网络首发:2015-02-10,
纸质出版:2015
移动端阅览
崔继岳, 梅魁志, 刘冬冬, 等. 面向OpenCL的Mali GPU仿真器构建研究[J]. 西安交通大学学报, 2015,49(2):20-24+68.
Construction of Embedded Mali GPU Simulator for OpenCL[J]. 2015, 49(2): 20-24+68.
崔继岳, 梅魁志, 刘冬冬, 等. 面向OpenCL的Mali GPU仿真器构建研究[J]. 西安交通大学学报, 2015,49(2):20-24+68. DOI: 10.7652/xjtuxb201502004.
Construction of Embedded Mali GPU Simulator for OpenCL[J]. 2015, 49(2): 20-24+68. DOI: 10.7652/xjtuxb201502004.
针对嵌入式GPU通用计算的仿真器构建需求
通过对通用图形处理单元仿真器(general purpose graphics processing unit-simulator
GPGPU-sim)的计算核心、存储结构与Mali GPU的异同进行比较分析
首先建立面向OpenCL的Mali GPU仿真器的流程与结构
并设计计算单元数、寄存器数、最小并行粒度等GPU微体系结构参数的获取方法
在对GPGPU-sim进行修改和配置后
实现了对特定GPU架构的仿真器构建。使用矩阵相乘、图像处理等OpenCL程序对仿真器的准确性进行测试
以程序在仿真器和硬件平台上的执行周期数差距作为评估依据。实验结果表明:对于测试程序集中优化前的OpenCL程序
其中70%的程序在两个平台上的运行周期数差距不超过30%; 对于优化后的OpenCL程序
其中90%的程序的运行周期数差距不超过30%。由此证明
构建的GPU仿真器能够满足OpenCL程序的仿真与性能评估。
The similarities and differences between GPGPU-sim and Mali GPU in computing cores and the storage structure are analyzed and compared
and simulating procedures and structures of Mali GPUs for OpenCL are built up to develop simulators for the general-purpose computing on embedded GPU. Methods to obtain the GPU microarchitecture parameters such as the computing unit number
the number of registers and the minimum parallel granularity are designed
and then the GPGPU-sim is configured and modified to construct specific GPU simulators. The accuracy of the simulator is tested through comparisons of running OpenCL programs
such as matrix multiplication and image processing on a real GPU and the simulator
and the difference between running cycles on the real GPU and the simulator is used as evaluation. Results show that the cycle differences are within 30% for about 70% OpenCL programs with simple implementation
and the cycle differences are within 30% for about 90% OpenCL programs with optimization. Therefore
it can be concluded that the constructed simulator meets the requirements of simulating and evaluating OpenCL programs on the embedded GPU.
NVIDIA. NVIDIA GeForce 8800 GPU architecture overview, TB-02787-001_V01[R]. Santa Clara, CA, USA: NVIDIA Corporation, 2006.
BAKHODA A, YUAN G L, FUNG W W L, et al. Analyzing CUDA workloads using a detailed GPU simulator [C]∥Proceedings of the IEEE International Symposium on Performance Analysis of Systems and Software. Piscataway, NJ, USA: IEEE, 2009: 163-174.
AAMODT T M, FUNG W W L, SINGH I, et al. GPGPU-Sim 3.x manual[EB/OL].(2012-08-08)[2013-08-08]. http:∥gpgpu-sim.org/manual/index. php/GPGPU-Sim_3.x_Manual.
WONG H, PAPADOPOULOU M M, SADOOGHI-ALVANDI M, et al. Demystifying GPU microarchitecture through microbenchmarking [C]∥Proceedings of the IEEE International Symposium on Performance Analysis of Systems and Software. Piscataway, NJ, USA: IEEE, 2010: 235-246.
TAYLOR R, LI Xiaoming. A micro-benchmark suite for AMD GPUs [C]∥Proceedings of the 39th International Conference on Parallel Processing Workshops. Washington, DC, USA: IEEE Computer Society, 2010: 387-396.
杨海燕, 史晓华, 孙清越, 等. 面向OpenCL的GPGPU微基准测试程序集的研究与实现 [J]. 系统工程与电子技术, 2013, 35(12): 2631-2642.YANG Haiyan, SHI Xiaohua, SUN Qingyue, et al. OpenCL micro benchmarks: testing the performance of GPGPU software and hardware architecture [J]. Systems Engineering and Electronics, 2013, 35(12): 2631-2642.
丑文龙,梅魁志,高增辉,等.ARM GPU的多任务调度设计与实现.2014,48(12):87-92.[doi:10.7652/xjtuxb2014120 14]
张虹,郑霄,赵丹.GPU加速窦房结计算机仿真的实现及优化.2014,48(7):60-64.[doi:10.7652/xjtuxb201407011]
李亮,王恩东,朱正东,等.ARM GPU的多任务调度设计与实现.2013,47(10):44-50.[doi:10.7652/xjtuxb201310008]
张保,曹海军,董小社,等.面向图形处理器重叠通信与计算的数据划分方法.2011,45(4):1-4.[doi:10.7652/xjtuxb2011 04001]
0
浏览量
4
下载量
1
CSCD
关联资源
相关文章
相关作者
相关机构
京公网安备11010802024621