西北工业大学计算机学院,西安,710072
网络首发:2010-10-10,
纸质出版:2010
移动端阅览
王得利, 高德远. 片上多核中一种共享感知的数据主动推送Cache技术[J]. 西安交通大学学报, 2010,44(10):18-23.
A Sharing-Aware Active Pushing Cache Technology on Chip-Multiprocessor[J]. 2010, 44(10): 18-23.
针对片上多核处理器的二级Cache访问延时持续增加以及并行程序在运行时线程间执行速率差异大的问题
提出了一种基于共享感知的数据主动推送Cache技术(SAAPC).SAAPC技术充分考虑并行程序的系统性能由速度最慢的线程所决定这一重要特性
根据并行线程间读数据共享程度高以及共享读数据访问局部性好的特征
采用基于指令的方法来预测共享读数据流
在后行线程需要共享数据之前将其主动推送至该线程的一级Cache中去
从而减少较慢线程的数据访问延时
提高执行速率
降低较慢线程与先行线程间执行速率的差异.SAAPC技术避免了预取技术所带来的额外片外带宽增加的缺点.使用SESC模拟器对来自于SPLASH2测试程序集的5个存储敏感型并行程序进行了测试仿真
结果表明
与传统的共享Cache相比
使用SAAPC技术减少了并行线程间执行速率的差异
系统的每周期指令数平均提高了7%
最高达到13.1%.
A sharing-aware active pushing Cache technology(SAAPC)is proposed to solve the problems of the increasing L2 Cache latency and the high deviation of progressive rates among the simultaneous threads in parallel applications when they are running on chip multi-processors. SAAPC fully takes the important characteristic into consideration that the whole system performances of parallel applications are constrained by the slowest thread in parallel phases. Based on the high share degree of read Cache blocks among different threads and the locality of the shared read accesses
SAPPC exploits the program counter to predict the shared data streams. The shared data are actively pushed to the L1 Caches of slower threads before it is needed so that the data access latencies for the slower threads are reduced and the progressive rates are increased. Therefore
the deviation of progressive rates is decreased. SAAPC avoids the problem caused by increasing off-chip bandwidth demand of the prefetch technique due to its inaccuracy. 5 memory intensive parallel programs from SPLASH2 benchmark suit are simulated using the simulator called SESC. Experimental results and comparisons with conventional shared Cache show that the SAAPC reduces the progressive rates deviation
and the average system instruction per cycle improvement is 7% and can be up to 13.1%.
KUNLE O, BASEM A, LANCER H. et al. The case for a single-chip multiprocessor [J]. ACM Sigplan Notices, 1996, 31(9):2-11.
CHEN Yu, LI Wenlong, LIN Junmin, et al. Data sharing analysis of emerging parallel media mining workloads [C]∥Proceedings of 2008 High Performance Computing. Berlin, Germany: Springer, 2008: 87-96.
CHEN Yu, LI Wenlong, KIM C K, et al. Efficient shared Cache management through sharing-aware replacement and streaming-aware insertion policy[C]∥Proceedings of the 2009 IEEE International Symposium on Parallel Distributed Processing. Los Alamitos, CA, USA: IEEE Computer Society, 2009:1-11.
BRADFORD M B, DAVID A W. Managing wire delay in large chip-multiprocessor Caches [C]∥Proceedings of the 37th annual IEEE/ACM International Symposium on Microarchitecture. Los Alamitos, CA, USA:IEEE Computer Society,2004:319-330.
STEVEN C W,MORIYOSHI O.The SPLASH-2 programs: characterization and methodological considerations [C]∥Proceedings of the 22nd Annual International Symposium on Computer Architecture. New York, USA: ACM, 1995:24-36.
THOMAS F W, STEPHEN S, NIKOLAOS H, et al. Temporal streaming of shared memory [C]∥ Proceedings of the 32nd Annual International Symposium on Computer Architecture. Los Alamitos, CA, USA: IEEE Computer Society, 2005:222-233.
STEFANOS K, JAMES R G. Improving CC-NUMA performance using instruction-based prediction[C]∥Proceedings of the 5th International Symposium on High Performance Computer Architecture. Los Alamitos, CA, USA: IEEE Computer Society, 1999:161-170.
0
浏览量
4
下载量
0
CSCD
关联资源
相关文章
相关作者
相关机构
京公网安备11010802024621