A hybrid deduplication method(Hy-Dedup)is adopted to solve the problem that the deduplication efficiency in the cloud storage system is not high for traditional mode online/offline deduplication
and the method performs effective data deduplication by combining online and offline modes. This method clusters fingerprint indices according the type of loads in online deduplication stage by adopting the fingerprint caching technology. The temporal local consistency of the duplicated data in data stream is estimated and the spatial local consistency is evaluated by setting different deduplication thresholds to reduce the disk fragments. The problem that the cache cannot be hit because lack of local consistency in the offline deduplication phase will be solved. The duplicated data is significantly reduced by this method while maintaining the I/O performance and the system throughput. Experimental results and a comparison with iDedup show that Hy-Dedup improves the online deduplication ratio by up to 35.9% and the disk capacity requirement reduces by 41.36%. It is concluded that the proposed method can achieve high-deciding deduplication in the cloud storage system
WANG Longxiang, ZHANG Xingjun, ZHU Guofeng, et al. A grouping prediction method based on undirected graph traversal in de-duplication system [J]. Journal of Xi'an Jiaotong University, 2013, 47(10): 51-56.
NISHA T R, ABIRAMI S, MANOHAR E. Experimental study on chunking algorithms of data deduplication system on large scale data [C]∥Proceedings of the International Conference on Soft Computing Systems, Advances in Intelligent Systems and Computing. Berlin, Germany: Springer, 2016: 91-98.
FU Yinjin, XIAO Nong, LIU Fang, et al. Deduplication based storage optimization technique for virtual desktop [J]. Journal of Computer Research and Development, 2012, 49(S1): 125-130.
SUDHAKARAN S, TREESA M. A survey on data deduplication in large scale data [J]. International Journal of Computer Applications, 2017, 165(1): 1-4.
XIA W, JIANG H, FENG D, et al. Similarity and locality based indexing for high performance data deduplication [J]. IEEE Transactions on Computers, 2015 64(4): 1162-1176.
SRINIVASAN K, BISSON T, GOODSON G, et al. iDedup: latency-aware, inline data deduplication for primary storage [C]∥Proceedings of the USENIX Conference on File and Storage Technologies. New York, USA: USENIX, 2012: 24.
MAO Bo, JIANG Hong, WU Suzhen, et al. POD: performance oriented I/O deduplication for primary storage systems in the cloud [C]∥Proceedings of the 2014 IEEE International Parallel and Distributed Processing Symposium. Piscataway, NJ, USA: IEEE, 2014: 767-776.
KAISER J, SÜß T, NAGEL L, et al. Sorted deduplication: how to process thousands of backup streams [C]∥ Proceedings of the 2016 32th Symposium on Mass Storage Systems and Technologies. Piscataway, NJ, USA: IEEE, 2016: 07897082.
FU Yinjin, JIANG Hong, XIAO Nong, et al. AA-dedupe: an application-aware source deduplication approach for cloud backup services in the personal computing environment [C]∥Proceedings of the IEEE International Conference on Cluster Computing. Piscataway, NJ, USA: IEEE, 2011: 112-120.
LI Wenji, JEAN-BAPTISE G, RIVEROS J, et al. Cachededup: in-line deduplication for flash caching [C]∥ Proceedings of the 14th USENIX Conference on File and Storage Technologies. New York, USA: USENIX, 2016: 301-314.
PARK D, RAN Ziqi, NAM Y J, et al. A lookahead read cache: improving read performance for deduplication backup storage [J]. Journal of Computer Science and Technology, 2017, 32(1): 26-40.
VITTER J S. Random sampling with a reservoir [J]. ACM Transactions on Mathematical Software, 1985, 11(1): 37-57.
VALIANT G, VALIANT P. Estimating the unseen: improved estimators for entropy and other properties [J]. Advances in Neural Information Processing Systems, 2013, 64(6): 2157-2165.
LIU Chuanyi, LU Yingping, SHI Chunhui, et al. ADMAD: application-driven metadata aware de-duplication archival storage system [C]∥Proceedings of the 2008 5th IEEE International Workshop on Storage Network Architecture and Parallel. Piscataway, NJ, USA: IEEE, 2008: 29-35.
NG C H, LEE P P C. RevDedup: a reverse deduplication storage system optimized for reads to latest backups [C]∥Proceedings of the 4th Asia-Pacific Workshop on Systems. New York, USA, ACM, 2013: 1-7.
LI Wenji, JEAN-BAPTISE G, RIVEROS J, et al. Cachededup: in-line deduplication for flash caching [C]∥Proceedings of the 14th USENIX Conference on File and Storage Technologies. New York USA: USENIX, 2016: 301-314.
MANDAL S, KUENNING G, OK D, et al. Using hints to improve inline block-layer deduplication [C]∥Proceedings of the 14th USENIX Conference on File and Storage Technologies. New York, USA: USENIX, 2016: 315-322.
FAN Ziqi, WU Fenggang, PARK D, et al. Hibachi: a cooperative hybrid cache with NVRAM and DRAM for storage arrays [C]∥Proceedings of the 2017 33th IEEE Symposium on Mass Storage Systems and Technologies. Piscataway, NJ, USA: IEEE, 2017: 1-12.