1. 清华大学智能与网络化系统研究中心,北京,100084
2. 国家计算机网络应急技术处理协调中心,北京,100029
网络首发:2015-12-10,
纸质出版:2015
移动端阅览
孙立远 1, 2, 周亚东 3, 等. 利用信息传播特性的中文网络新词发现方法[J]. 西安交通大学学报, 2015,49(12):59-64.
A Method of Discovering New Chinese Words from Internet Based on Information Propagation[J]. 2015, 49(12): 59-64.
孙立远 1, 2, 周亚东 3, 等. 利用信息传播特性的中文网络新词发现方法[J]. 西安交通大学学报, 2015,49(12):59-64. DOI: 10.7652/xjtuxb201512010.
A Method of Discovering New Chinese Words from Internet Based on Information Propagation[J]. 2015, 49(12): 59-64. DOI: 10.7652/xjtuxb201512010.
针对已有方法识别出的网络中文新词生命周期短且很快不再为人们所用的问题
提出了一种基于信息传播特性的中文新词发现方法。该方法结合“新词传播范围广、持续时间长”的特点
从用户覆盖率、话题覆盖率和新词生命周期3个方面设计统计量; 采用N-gram算法得到候选词串列表; 用基于词频和词语灵活度的方法过滤垃圾词串。实验中以微博文本作为语料来源
与已有方法相比
用户特性使得新词识别的准确率提高了11%
话题特性使准确率提高了10%
时间特性使准确率提高了13%
综合用户、话题和时间的方法使准确率提高了16%。实验结果表明:该方法中的每个特性都提高了中文网络新词识别的准确率
而且同时考虑3种特性的准确率比只考虑单一特性的高。
A method of discovering new Chinese words from Internet based on information propagation is proposed to solve the problems that the recognizing results of existing methods always have short life cycles and will not be used again in soon. The method combines the characteristics of new words such as widely spreading and long lasting
and three statistics
i.e. coverage rate of users
coverage rate of topics and life cycle of a new word
are defined. The N-gram algorithm is applied to generate candidates of new words
then the word candidates are filtered bade on word frequency and word flexibility. Experiments with the text of microblogs as corpus and comparisons with the existing methods show that the user statistic enhances the accuracy rate of recognizing new words by 11%
the topic statistic enhances the accuracy rate by 10%
and the time statistic enhances the accuracy rate by 13%. When the three statistics are combined
the accuracy rate is raised by 16%. It can be concluded that each single statistic considered by the proposed method can enhance the accuracy rate
and more accurate rate can be obtained by considering the combination of the three statistics rather than just considering one statistic.
张海军, 史树敏, 朱朝勇, 等. 中文新词识别技术综述 [J]. 计算机科学, 2010, 37(3): 6-10.
ZHANG Haijun, SHI Shumin, ZHU Zhaoyong, et al. Survey of Chinese new words identification [J]. Computer Science, 2010, 37(3): 6-10.
霍帅, 张敏, 刘奕群, 等. 基于微博内容的新词发现方法 [J]. 模式识别与人工智能, 2014, 27(2): 141-145.
HUO Shuai, ZHANG Min, LIU Yiqun, et al. New word discovery in microblog content [J]. Pattern Recognition and Artificial Intelligence, 2014, 27(2): 141-145.
苏其龙. 微博新词发现研究 [D]. 哈尔滨: 哈尔滨工业大学, 2013.
杨辉. 汉语新词语发现及其词性标注方法研究 [D]. 上海: 复旦大学, 2008.
邹纲, 刘洋, 刘群, 等. 面向Internet的中文新词语检测 [J]. 中文信息学报, 2004, 18(6): 1-9.
ZOU Gang, LIU Yang, LIU Qun, et al. Internet-oriented Chinese new words detection [J]. Journal of Chinese Information Processing, 2004, 18(6): 1-9.
SUI Zhifang, CHEN Yirong. The research on the automatic term extraction in the domain of information science and technology [C]∥Proceedings of the 5th East Asia Forum of Terminology. Beijing, China: China National Institute of Standardization, 2002: 17-21.
HIDEKI I. Japanese named entity recognition based on a simple rule generator and decision tree learning [C]∥Proceedings of the 39th Annual Meeting on Association for Computational Linguistics. Stroudsburg, PA, USA: Association for Computational Linguistics, 2001: 314-321.
罗盛芬, 孙茂松. 基于字串内部结合紧密度的汉语自动抽词实验研究 [J]. 中文信息学报, 2003, 17(3): 9-14.
LUO Shengfen, SUN Maosong. Chinese word extraction based on the internal associative strength of character strings [J]. Journal of Chinese Information Processing, 2003, 17(3): 9-14.
YE Yunming, WU Qingyao, LI Yan, et al. Unknown Chinese word extraction based on variety of overlapping strings [J]. Information Processing and Management, 2013, 49(2): 497-512.
HUANG J H, POWERS D. Chinese word segmentation based on contextual entropy [C]∥Proceedings of the 17th Asian Pacific Conference on Language, Information and Computation. Piscataway, NJ, USA: IEEE, 2003: 152-158.
孙立远, 袁睿翕, 卞小丁. 一种中文网页新词自动获取方法: 中国, ZL 200910237979.3 [P]. 2011-06-01.
周亚东. 在线社会网络热点话题识别与动态传播建模与分析研究 [D]. 西安: 西安交通大学, 2011.
李文坤, 张仰森, 陈若愚. 基于词内部结合度和边界自由度的新词发现 [J]. 计算机应用研究, 2015, 32(8): 51-55.
LI Wenkun, ZHANG Yangsen, CHEN Ruoyu. New word detection based on inner combination degree and boundary freedom degree of word [J]. Application Research of Computers, 2015, 32(8): 51-55.
杨攀,桂小林,安健,等.利用贝叶斯原理在隐私保护数据上进行分类的方法.2015,49(4):46-52.[doi:10.7652/xjtuxb 201504008]
李刘强,桂小林,安健,等.采用模糊层次聚类的社会网络重叠社区检测算法.2015,49(2):6-13.[doi:10.7652/xjtuxb 201502002]
李长路,王劲林,郭志川,等.两阶段密度意识子空间聚类模型.2014,48(10):108-114.[doi:10.7652/xjtuxb201410017]
李涛,肖南峰.应用相似度测量的图离群点检测方法.2014,48(8):67-72.[doi:10.7652/xjtuxb201408012]
陈家旭,唐亚哲,胡成臣,等.延迟容忍网络中基于地点偏好的社会感知多播路由协议设计.2014,48(6):13-18.[doi:10.7652/xjtuxb201406003]
张赛,徐恪,李海涛.微博类社交网络中信息传播的测量与分析.2013,47(2):124-130.[doi:10.7652/xjtuxb201302021]
莫同,褚伟杰,李伟平,等.采用超图的微博群落感知方法.2012,46(11):120-126.[doi:10.7652/xjtuxb201211022]
豆增发,高琳.利用膜粒子群优化和信息熵的医学文本特征选择.2012,46(4):45-51.[doi:10.7652/xjtuxb201204008]
陈刚,蔡远利,穆静,等.海量信息异常检测问题的异常概率排序算法.2011,45(4):36-40.[doi:10.7652/xjtuxb201104 007]
刘京鑫,孙剑,孟德宇.基于视觉原理的分类算法.2010,44(10):116-119.[doi:10.7652/xjtuxb201010022]
冯少荣,张东站.高效的用户访问预测新算法.2010,44(4):28-33.[doi:10.7652/xjtuxb201004007]
李小虎,杜海峰,庄健,等.基于小世界原理的模型降阶优化研究.2009,43(1):108-113.[doi:10.7652/xjtuxb200901024]
朱虎明,焦李成.基于免疫记忆克隆的特征选择.2008,42(6):679-682.[doi:10.7652/xjtuxb200806007]
周亚东,孙钦东,管晓宏,等.流量内容词语相关度的网络热点话题提取.2007,41(10):1142-1145.[doi:10.7652/xjtuxb 200710004]
杜海峰,李树茁,Marcus W.Feldman,等.基于先验知识与模块性的网络社区结构探测算法.2007,41(6):750-754.[doi:10.7652/xjtuxb200706026]
0
浏览量
4
下载量
2
CSCD
关联资源
相关文章
相关作者
相关机构
京公网安备11010802024621