An algorithm for Chinese semantic orientation calculation that uses distribution similarity is proposed to solve the problem that existing methods take less implied semantic into consideration in semantic orientation inference. The Chinese semantic orientation calculation is carried out in two steps. The first step calculates the distribution similarities using dependency grammar analysis and statistical tools. HowNet and Chinese conjunction features are introduced in semantic similarity calculation to optimize the corpus-based statistical results. The second step adopts an undirected weighted graph clustering algorithm to infer semantic orientation.Because it is an NP-hard problem to obtain the optimal clustering solution
a greedy algorithm is used to get an approximate solution. Experiments on the testing corpus show that the accuracy of the proposed algorithm is 80% and is better than both the corpus-based statistic algorithm and the HowNet-based algorithm. The results demonstrate that the proposed method is feasible and effective to improve the accuracy of Chinese semantic orientation calculation.
关键词
Keywords
references
HATZIVASSILOGLOU V, MCKEOWN K R. Predicting the semantic orientation of adjectives [C]∥Proceedings of the 35th Annual Meeting of the ACL and the 8th Conference of the European Chapter of the ACL. Stroudsburg, PA, USA: Association for Computational Linguistics, 1997: 174-181.
KAMPS J, MARX M, MOKKEN R J, et al. Using WordNet to measure semantic orientations of adjectives [C]∥Proceedings of the 4th International Conference on Language Resources and Evaluation. Paris, France: European Language Resources Association, 2004: 1115-1118.
ZHU Yanlan, MIN Jin, ZHOU Yaqian, et al. Semantic orientation computing based on HowNet [J]. Journal of Chinese Information Processing, 2006, 20(1):14-20.
LIN D. Automatic retrieval and clustering of similar words [C]∥Proceedings of COLING/ACL.Stroudsburg, PA, USA: Association for Computational Linguistics, 1998: 768-774.
MCCARTHY D, KOELING R, WEEDS J, et al. Finding predominant word senses in untagged text [C]∥Proceedings of the 42nd Annual Meeting of the Association for Computational Linguistics. Stroudsburg, PA, USA: Association for Computational Linguistics, 2004:279-286.
ZHOU Junsheng, HUANG Shujian, CHEN Jiajun, et al. A new graph clustering algorithm for Chinese noun phrase coreference resolution [J]. Journal of Chinese Information Processing, 2007, 21(2): 77-82.