A sharing-aware active pushing Cache technology(SAAPC)is proposed to solve the problems of the increasing L2 Cache latency and the high deviation of progressive rates among the simultaneous threads in parallel applications when they are running on chip multi-processors. SAAPC fully takes the important characteristic into consideration that the whole system performances of parallel applications are constrained by the slowest thread in parallel phases. Based on the high share degree of read Cache blocks among different threads and the locality of the shared read accesses
SAPPC exploits the program counter to predict the shared data streams. The shared data are actively pushed to the L1 Caches of slower threads before it is needed so that the data access latencies for the slower threads are reduced and the progressive rates are increased. Therefore
the deviation of progressive rates is decreased. SAAPC avoids the problem caused by increasing off-chip bandwidth demand of the prefetch technique due to its inaccuracy. 5 memory intensive parallel programs from SPLASH2 benchmark suit are simulated using the simulator called SESC. Experimental results and comparisons with conventional shared Cache show that the SAAPC reduces the progressive rates deviation
and the average system instruction per cycle improvement is 7% and can be up to 13.1%.
关键词
Keywords
references
KUNLE O, BASEM A, LANCER H. et al. The case for a single-chip multiprocessor [J]. ACM Sigplan Notices, 1996, 31(9):2-11.
CHEN Yu, LI Wenlong, LIN Junmin, et al. Data sharing analysis of emerging parallel media mining workloads [C]∥Proceedings of 2008 High Performance Computing. Berlin, Germany: Springer, 2008: 87-96.
CHEN Yu, LI Wenlong, KIM C K, et al. Efficient shared Cache management through sharing-aware replacement and streaming-aware insertion policy[C]∥Proceedings of the 2009 IEEE International Symposium on Parallel Distributed Processing. Los Alamitos, CA, USA: IEEE Computer Society, 2009:1-11.
BRADFORD M B, DAVID A W. Managing wire delay in large chip-multiprocessor Caches [C]∥Proceedings of the 37th annual IEEE/ACM International Symposium on Microarchitecture. Los Alamitos, CA, USA:IEEE Computer Society,2004:319-330.
STEVEN C W,MORIYOSHI O.The SPLASH-2 programs: characterization and methodological considerations [C]∥Proceedings of the 22nd Annual International Symposium on Computer Architecture. New York, USA: ACM, 1995:24-36.
THOMAS F W, STEPHEN S, NIKOLAOS H, et al. Temporal streaming of shared memory [C]∥ Proceedings of the 32nd Annual International Symposium on Computer Architecture. Los Alamitos, CA, USA: IEEE Computer Society, 2005:222-233.
STEFANOS K, JAMES R G. Improving CC-NUMA performance using instruction-based prediction[C]∥Proceedings of the 5th International Symposium on High Performance Computer Architecture. Los Alamitos, CA, USA: IEEE Computer Society, 1999:161-170.