An index table grouping prediction method is proposed based on undirected graph traversal to solve the problem that the data storage capacity of the de-duplication system is limited by the memory and is difficult to expand into large-scale. The method saves the index in a disk and maintains the cache of index in memory to expand the maximum storage capacity of system. The hit rate of the grouping prediction and the system performance are improved by grouping index entries based on undirected graph traversal. The method sets up and analyzes the undirected graph based on the features of data chunk sequences
and the groups generated by analyzing the graph is used in cache replacement. Experimental results show that since the proposed method bases on the cache prefetching and the hash table grouping
the index table cache hit rate of the method increases from 47% to 87.6%
and the maximum storage capacity of IDSMS system is 7.5 times higher than that of the existing method for the cache consuming only 10% of the index table size.
关键词
Keywords
references
DUBOIS L, AMALDAS M, SHEPPARD E. Key considerations as deduplication evolves into primary storage, White Paper 223310 [R]. Framingham, MA, USA: IDC, 2011.
GU Yu, LIU Chuanyi, SUN Linchun, et al. Reliability provision mechanism for large-scale de-duplication storage systems [J]. Journal of Tsinghua University: Science and Technology, 2010, 50(5): 739-744.
FU Yinjin, XIAO Nong, LIU Fang, et al. Deduplication based storage optimization technique for virtual desktop [J]. Journal of Computer Research and Development, 2012, 49(Supp 1): 125-130.
QUINLAN S, DORWARD S. Venti: a new approach to archival storage [C]∥Proceedings of the FAST 2002 Conference on File and Storage Technologies. Berkeley, CA, USA: USENIX, 2002: 89-101.
ZHU B, LI Kai, PATTERSON H. Avoiding the disk bottleneck in the data domain deduplication file system [C]∥Proceedings of the 6th USENIX Conference on File and Storage Technologies. Berkeley, CA, USA: USENIX, 2008: 269-282.
BHAGWAT D, ESHGHI K, LONG D D E, et al. Extreme binning: scalable, parallel deduplication for chunk-based file backup [C]∥Proceedings of the IEEE International Symposium on Modeling, Analysis and Simulation of Computer and Telecommunication Systems. Piscataway, USA: IEEE, 2009: 1-9.
LILLIBRIDGE M, ESHGHI K, BHAGWAT D, et al. Sparse indexing: large scale, inline deduplication using sampling and locality [C]∥Proceedings of the 7th Conference on File and Storage Technologies. Berkeley, CA, USA: USENIX, 2009: 111-123.
MIN J, YOON D, WON Y. Efficient deduplication techniques for modern backup operation [J]. IEEE Transactions on Computers, 2011, 60(6): 824-840.
BOBBARJUNG D R, JAGANNATHAN S, DUBNICKI C. Improving duplicate elimination in storage systems [J]. ACM Transactions on Storage, 2006, 2(4): 424-448.
TSUCHIYA Y, WATANABE T. DBLK: deduplication for primary block storage [C]∥Proceedings of the 27th IEEE Symposium on Mass Storage Systems and Technologies. Piscataway, USA: IEEE, 2011: 1-5.
TOMAZIC S, PAVLOVIC V, MILOVANOVIC J, et al. Fast file existence checking in archiving systems [J]. ACM Transactions on Storage, 2011, 7(1): 24-44.
WILDANI A, MILLER E L, RODEH O. Hands: a heuristically arranged non-backup in-line deduplication system, UCSC-SSRC-12-03 [R]. Piscataway, USA: IEEE, 2012.
WILDANI A, MILLER E, WARD L. Efficiently identifying working sets in block I/O streams [C]∥Proceedings of the 4th Annual International Conference on Systems and Storage. New York, USA: ACM, 2011: 46-57.
ZHU Guofeng, ZHANG Xingjun, WANG Longxiang, et al. An intelligent data de-duplication based backup system [C]∥Proceedings of the 15th International Conference on Network-Based Information Systems. Piscataway, USA: IEEE, 2012: 771-776.