A user identification algorithm based on blocking and bipartite graph is proposed to solve the problem of schema inconsistency in user identification across social networks. The schema alignment is avoided by extending the precise attributes matching in traditional blocking to approximately matching without using schema information. The weighted bipartite graph and the Kuhn Munkres(KM)algorithm are employed to calculate the similarity between the source and the candidate user profiles so that the problems of different attribute numbers as well as the semantic and syntactic heterogeneity between user profiles are solved. Public data of 965 users are collected from the Profilactic website and the proposed algorithm is evaluated using the recall
precision and F-measure metrics. Experimental results show that the proposed algorithm can implement cross-system and instance-level user identification without the use of schema information
and a comparison with the method using precise attributes matching shows that the recall of the algorithm is improved by 6.2%-9.5% and the F-measure is improved by 3%-4.2%.
关键词
Keywords
references
BODHIT A, AMIN K. Possible solutions of new user or item cold-start problem [J]. International Journal of Mathematics and Computer Research, 2013, 1(3): 123-128.
CARMAGNOLA F, CENA F. User identification for cross-system personalisation [J]. Information Sciences, 2009, 179(1): 16-32.
VOSECKY J, HONG D, SHEN V Y. User identification across multiple social networks [C]∥Proceedings of the First International Conference on Networked Digital Technologies. Piscataway, NJ, USA: IEEE, 2009: 360-365.
IOFCIU T, FANKHOUSER P, ABEL F, et al. Identifying users across social tagging systems [C]∥Proceedings of the 5th International AAAI Conference on Weblogs and Social Media. Palo Alto, California, USA: AAAI, 2011: 1-4.
RAAD E, CHBEIR R, DIPANDA A. User profile matching in social networks [C]∥Proceedings of the 13th International Conference on Network-Based Information Systems. Piscataway, NJ, USA: IEEE, 2010: 297-304.
MARTINEZ -VILLASENOR M L G, GONZALEZ-MENDOZA M. Process of concept alignment for interoperability between heterogeneous sources [C]∥Proceedings of the 11th Mexican International Conference on Advances in Artificial Intelligence. Berlin, Germany: Springer-Verlag, 2013: 311-320.
BAXTER R, CHRISTEN P, CHURCHES T. A comparison of fast blocking methods for record linkage [C]∥Proceedings of the First Workshop on Data Cleaning, Record Linkage and Object Consolidation. New York, USA: ACM, 2003: 25-27.
PAPADAKIS G, IOANNOU E, NIEDERÉE C, et al. Efficient entity resolution for large heterogeneous information spaces [C]∥Proceedings of the 4th ACM International Conference on Web Search and Data Mining. New York, USA: ACM, 2011: 535-544.
WHITE S. How to strike a match [EB/OL].(2010-05-03)[2012-05-25]. http:∥www.catalysoft.com/articles/StrikeAMatch.html.
LI Mohan, WANG Hongzhi, LI Jianzhong, et al.Duplicate record detection method based on optimal bipartite graph matching [J]. Journal of Computer Research and Development, 2009, 46(S2): 339-345.
CARMAGNOLA F, OSBORNE F, TORRE I. User data distributed on the social web: how to identify users on different social systems and collecting data about them [C]∥Proceedings of the First International Workshop on Information Heterogeneity and Fusion in Recommender Systems. New York, USA: ACM, 2010: 9-15.