1. 西安交通大学电子与信息工程学院,西安,710049
2. 西安建筑科技大学信息与控制工程学院,西安,710055
网络首发:2013-12-10,
纸质出版:2013
移动端阅览
叶娜 1, 2, 赵银亮 1, 等. 模式无关的社交网络用户识别算法[J]. 西安交通大学学报, 2013,47(12):19-25.
A Schema-Independent User Identification Algorithm in Social Networks[J]. 2013, 47(12): 19-25.
叶娜 1, 2, 赵银亮 1, 等. 模式无关的社交网络用户识别算法[J]. 西安交通大学学报, 2013,47(12):19-25. DOI: 10.7652/xjtuxb201312004.
A Schema-Independent User Identification Algorithm in Social Networks[J]. 2013, 47(12): 19-25. DOI: 10.7652/xjtuxb201312004.
针对识别社交网络用户时存在的模式不一致问题
提出了基于分块和二部图的用户识别算法。该算法通过将传统分块算法中的属性值精确匹配扩展为无模式信息下的属性值近似匹配
避免了传统用户识别时所需的模式对齐; 使用加权二部图及Kuhn Munkres(KM)最大权匹配算法进行源用户档案与待匹配用户档案间的相似度计算
解决了用户档案间属性个数不同及语义语法异构的问题。在社交网站Profilactic上采集了965个用户的公开数据
采用召回率、精确率和综合指标等评价指标对算法进行了实验评估。实验结果表明
所提算法能够不依赖模式信息进行实例级跨系统用户识别
与基于属性值精确匹配的算法相比
所提算法的召回率提高了6.2%~9.5%
综合评价指标提高了3%~4.2%。
A user identification algorithm based on blocking and bipartite graph is proposed to solve the problem of schema inconsistency in user identification across social networks. The schema alignment is avoided by extending the precise attributes matching in traditional blocking to approximately matching without using schema information. The weighted bipartite graph and the Kuhn Munkres(KM)algorithm are employed to calculate the similarity between the source and the candidate user profiles so that the problems of different attribute numbers as well as the semantic and syntactic heterogeneity between user profiles are solved. Public data of 965 users are collected from the Profilactic website and the proposed algorithm is evaluated using the recall
precision and F-measure metrics. Experimental results show that the proposed algorithm can implement cross-system and instance-level user identification without the use of schema information
and a comparison with the method using precise attributes matching shows that the recall of the algorithm is improved by 6.2%-9.5% and the F-measure is improved by 3%-4.2%.
BODHIT A, AMIN K. Possible solutions of new user or item cold-start problem [J]. International Journal of Mathematics and Computer Research, 2013, 1(3): 123-128.
CARMAGNOLA F, CENA F. User identification for cross-system personalisation [J]. Information Sciences, 2009, 179(1): 16-32.
VOSECKY J, HONG D, SHEN V Y. User identification across multiple social networks [C]∥Proceedings of the First International Conference on Networked Digital Technologies. Piscataway, NJ, USA: IEEE, 2009: 360-365.
IOFCIU T, FANKHOUSER P, ABEL F, et al. Identifying users across social tagging systems [C]∥Proceedings of the 5th International AAAI Conference on Weblogs and Social Media. Palo Alto, California, USA: AAAI, 2011: 1-4.
RAAD E, CHBEIR R, DIPANDA A. User profile matching in social networks [C]∥Proceedings of the 13th International Conference on Network-Based Information Systems. Piscataway, NJ, USA: IEEE, 2010: 297-304.
MARTINEZ -VILLASENOR M L G, GONZALEZ-MENDOZA M. Process of concept alignment for interoperability between heterogeneous sources [C]∥Proceedings of the 11th Mexican International Conference on Advances in Artificial Intelligence. Berlin, Germany: Springer-Verlag, 2013: 311-320.
BAXTER R, CHRISTEN P, CHURCHES T. A comparison of fast blocking methods for record linkage [C]∥Proceedings of the First Workshop on Data Cleaning, Record Linkage and Object Consolidation. New York, USA: ACM, 2003: 25-27.
PAPADAKIS G, IOANNOU E, NIEDERÉE C, et al. Efficient entity resolution for large heterogeneous information spaces [C]∥Proceedings of the 4th ACM International Conference on Web Search and Data Mining. New York, USA: ACM, 2011: 535-544.
WHITE S. How to strike a match [EB/OL].(2010-05-03)[2012-05-25]. http:∥www.catalysoft.com/articles/StrikeAMatch.html.
李默涵, 王宏志, 李建中, 等. 一种基于二分图最优匹配的重复记录检测算法 [J].计算机研究与发展, 2009, 46(S2): 339-345.
LI Mohan, WANG Hongzhi, LI Jianzhong, et al.Duplicate record detection method based on optimal bipartite graph matching [J]. Journal of Computer Research and Development, 2009, 46(S2): 339-345.
CARMAGNOLA F, OSBORNE F, TORRE I. User data distributed on the social web: how to identify users on different social systems and collecting data about them [C]∥Proceedings of the First International Workshop on Information Heterogeneity and Fusion in Recommender Systems. New York, USA: ACM, 2010: 9-15.
王龙翔,张兴军,朱国峰,等.重复数据删除中的无向图遍历分组预测方法.2013,47(10):51-56.[doi:10.7652/xjtuxb 201310009]
李清华,康海燕,苑晓姣,等.个性化搜索中用户兴趣模型匿名化研究.2013,47(4):131-136.[doi:10.7652/xjtuxb2013 04022]
张赛,徐恪,李海涛.微博类社交网络中信息传播的测量与分析.2013,47(2):124-130.[doi:10.7652/xjtuxb201302021]
陈国强,王宇平.采用离散粒子群算法的复杂网络重叠社团检测.2013,47(1):107-113.[doi:10.7652/xjtuxb201301021]
薛咏,冯博琴,刘卫涛.扩展主题图本体融合策略与算法.2011,45(10):13-18.[doi:10.7652/xjtuxb201110003]
豆增发,高琳.应用粒子群优化-条件随机域的文本生物实体识别.2010,44(12):38-42.[doi:10.7652/xjtuxb201012008]
鲁慧民,冯博琴,李旭.面向多源知识融合的扩展主题图相似性算法.2010,44(2):20-24.[doi:10.7652/xjtuxb201002005]
0
浏览量
4
下载量
2
CSCD
关联资源
相关文章
相关作者
相关机构
京公网安备11010802024621