Focusing on the problems that the text clustering has high time complexity
the algorithms that are independent on the distance are unsuitable for dynamic and changing corpus
and the multi-subject characteristics of a single text cannot be considered in traditional algorithms
an incremental algorithm of text soft clustering based on semantic sequence is proposed
in which the clustering candidate with minimum entropy overlap value is selected as a result cluster by using similarity relation of semantic sequences and calculating the coverage of similarity semantic sequences set. The dimensions of text vector space are decreased dramatically in the clustering procedure
so the computing time can be reduced. Since the semantic sequence is only related to text
it is available for incremental clustering. The comparison of experimental results shows that the algorithm can achieve higher precision than other algorithms under same conditions
especially for soft clustering of long texts set.
关键词
Keywords
references
Boley D, Gini M,Gross R,et al. Partitioning-based clustering for web document categorization [J]. Decision Support Systems, 1999, 27(3):329-341.
Zamir O, Etzioni O. Web document clustering: a feasibility demonstration [C]∥Proceedings of the 19th ACM SIGIR Conference on Research and Development in Information Retrieval. New York: ACM Press, 1998:46-54.
Dhillon I S, Guan Y,Kogan J. Co-clustering documents and words using bipartite spectral graph partitioning [C]∥Proceedings of the 7th ACM SIGKDD Conference on Knowledge Discovery and Data Mining. New York: ACM Press, 2001:269-274.
Bao J P, Shen J Y, Liu X D, et al. Semantic sequence kin: a method of document copy detection [C]∥Proceedings of the 8th Pacific-Asia Conference on Knowledge Discovery and Data Mining. Berlin:Springer-Verlag, 2004:529-538.
Beil F, Ester M, Xu X W. Frequent term-based text clustering [C]∥Proceedings of the 8th ACM SIGKDD Conference on Knowledge Discovery and Data Mining. New York: ACM Press, 2002:436-442.
Cover T M, Thomas J A. Elements of information theory [M]. New York: Wiley, 1991.
Slonim N, Tishby N. Document clustering using word clusters via the information bottleneck method [C]∥Proceedings of the 21st ACM SIGIR Conference on Research and Development in Information Retrieval. New York: ACM Press, 2000:208-215.