A method of discovering new Chinese words from Internet based on information propagation is proposed to solve the problems that the recognizing results of existing methods always have short life cycles and will not be used again in soon. The method combines the characteristics of new words such as widely spreading and long lasting
and three statistics
i.e. coverage rate of users
coverage rate of topics and life cycle of a new word
are defined. The N-gram algorithm is applied to generate candidates of new words
then the word candidates are filtered bade on word frequency and word flexibility. Experiments with the text of microblogs as corpus and comparisons with the existing methods show that the user statistic enhances the accuracy rate of recognizing new words by 11%
the topic statistic enhances the accuracy rate by 10%
and the time statistic enhances the accuracy rate by 13%. When the three statistics are combined
the accuracy rate is raised by 16%. It can be concluded that each single statistic considered by the proposed method can enhance the accuracy rate
and more accurate rate can be obtained by considering the combination of the three statistics rather than just considering one statistic.
HUO Shuai, ZHANG Min, LIU Yiqun, et al. New word discovery in microblog content [J]. Pattern Recognition and Artificial Intelligence, 2014, 27(2): 141-145.
ZOU Gang, LIU Yang, LIU Qun, et al. Internet-oriented Chinese new words detection [J]. Journal of Chinese Information Processing, 2004, 18(6): 1-9.
SUI Zhifang, CHEN Yirong. The research on the automatic term extraction in the domain of information science and technology [C]∥Proceedings of the 5th East Asia Forum of Terminology. Beijing, China: China National Institute of Standardization, 2002: 17-21.
HIDEKI I. Japanese named entity recognition based on a simple rule generator and decision tree learning [C]∥Proceedings of the 39th Annual Meeting on Association for Computational Linguistics. Stroudsburg, PA, USA: Association for Computational Linguistics, 2001: 314-321.
LUO Shengfen, SUN Maosong. Chinese word extraction based on the internal associative strength of character strings [J]. Journal of Chinese Information Processing, 2003, 17(3): 9-14.
YE Yunming, WU Qingyao, LI Yan, et al. Unknown Chinese word extraction based on variety of overlapping strings [J]. Information Processing and Management, 2013, 49(2): 497-512.
HUANG J H, POWERS D. Chinese word segmentation based on contextual entropy [C]∥Proceedings of the 17th Asian Pacific Conference on Language, Information and Computation. Piscataway, NJ, USA: IEEE, 2003: 152-158.
LI Wenkun, ZHANG Yangsen, CHEN Ruoyu. New word detection based on inner combination degree and boundary freedom degree of word [J]. Application Research of Computers, 2015, 32(8): 51-55.