西北工业大学计算机学院,西安,710072
网络首发:2009-06-10,
纸质出版:2009
移动端阅览
赵煜, 蔡皖东, 樊娜, 等. 利用词汇分布相似度的中文词汇语义倾向性计算[J]. 西安交通大学学报, 2009,43(6):33-37.
Computing Chinese Semantic Orientation Via Distributional Similarity[J]. 2009, 43(6): 33-37.
针对现有中文词汇语义倾向性计算方法存在较少考虑深层语义影响因素的问题
提出了一种利用词汇分布相似度的中文语义倾向性计算方法.该方法分2个步骤完成:①利用依存句法分析和统计工具获取词汇在语料库中的分布相似度
并综合知网(HowNet)和汉语连词特征信息优化语料库统计结果
计算中文词汇间的语义相似度; ②采用无向带权图划分的聚类方法来实现中文词汇语义倾向推断.由于获取最优聚类结果是一个NP难问题
所以采用贪心算法求解近似最优值.通过在自建的语料库上进行测试
并与利用语料库统计信息、利用HowNet等2个词汇语义倾向性计算系统进行比较
结果是所提方法的准确率达到了80%
表明在提高中文词汇语义倾向性计算的准确性方面是可行、有效的.
An algorithm for Chinese semantic orientation calculation that uses distribution similarity is proposed to solve the problem that existing methods take less implied semantic into consideration in semantic orientation inference. The Chinese semantic orientation calculation is carried out in two steps. The first step calculates the distribution similarities using dependency grammar analysis and statistical tools. HowNet and Chinese conjunction features are introduced in semantic similarity calculation to optimize the corpus-based statistical results. The second step adopts an undirected weighted graph clustering algorithm to infer semantic orientation.Because it is an NP-hard problem to obtain the optimal clustering solution
a greedy algorithm is used to get an approximate solution. Experiments on the testing corpus show that the accuracy of the proposed algorithm is 80% and is better than both the corpus-based statistic algorithm and the HowNet-based algorithm. The results demonstrate that the proposed method is feasible and effective to improve the accuracy of Chinese semantic orientation calculation.
HATZIVASSILOGLOU V, MCKEOWN K R. Predicting the semantic orientation of adjectives [C]∥Proceedings of the 35th Annual Meeting of the ACL and the 8th Conference of the European Chapter of the ACL. Stroudsburg, PA, USA: Association for Computational Linguistics, 1997: 174-181.
KAMPS J, MARX M, MOKKEN R J, et al. Using WordNet to measure semantic orientations of adjectives [C]∥Proceedings of the 4th International Conference on Language Resources and Evaluation. Paris, France: European Language Resources Association, 2004: 1115-1118.
苏祺. 面向问答系统的情感倾向分析研究 [D]. 北京: 北京大学信息科学技术学院, 2007.
朱嫣岚, 闵锦, 周雅倩,等. 基于HowNet的词汇语义倾向计算 [J]. 中文信息学报, 2006, 20(1): 14-20.
ZHU Yanlan, MIN Jin, ZHOU Yaqian, et al. Semantic orientation computing based on HowNet [J]. Journal of Chinese Information Processing, 2006, 20(1):14-20.
LIN D. Automatic retrieval and clustering of similar words [C]∥Proceedings of COLING/ACL.Stroudsburg, PA, USA: Association for Computational Linguistics, 1998: 768-774.
MCCARTHY D, KOELING R, WEEDS J, et al. Finding predominant word senses in untagged text [C]∥Proceedings of the 42nd Annual Meeting of the Association for Computational Linguistics. Stroudsburg, PA, USA: Association for Computational Linguistics, 2004:279-286.
刘群,李素建. 基于《知网》的词汇语义相似度计算[J]. 中文计算语言学期刊, 2002, 17(2): 59-76.
LIU Qun, LI Sujian. Word similarity computing based on HowNet[J]. Computational Linguistics and Chinese Language Processing, 2002, 17(2): 59-76.
NEWMAN M E J. Mixing patterns in networks [J]. Physical Review: E, 2003, 67(22):26126.
NEWMAN M E J, GIRVAN M. Finding and evaluating community structure in networks [J]. Physical Review: E, 2004, 69(22):26113.
周俊生,黄书剑,陈家骏,等. 一种基于图划分的无监督汉语指代消解算法[J]. 中文信息学报, 2007,21(2): 77-82.
ZHOU Junsheng, HUANG Shujian, CHEN Jiajun, et al. A new graph clustering algorithm for Chinese noun phrase coreference resolution [J]. Journal of Chinese Information Processing, 2007, 21(2): 77-82.
【本刊相关文献链接】
广义随机Petri网下的组合Web服务建模与评价. 2008, 42(8): 967-971.
可用性语义Web服务的通用发现机制. 2008, 42(6): 659-663.
结合受控词汇表的生物基因本体标注与分类. 2008, 42(2): 171-174.
八邻域网格聚类的多样性XML文档近似查询算法. 2007, 41(8): 907-911.
一种增量式文本软聚类算法. 2007, 41(4): 398-401.
Web流语义感知的改进队列管理算法. 2006, 40(10): 1047-1051.
0
浏览量
4
下载量
1
CSCD
关联资源
相关文章
相关作者
相关机构
京公网安备11010802024621