西安交通大学电子与信息工程学院,西安,710049
网络首发:2007-04-10,
纸质出版:2007
移动端阅览
冯中慧, 鲍军鹏, 沈钧毅. 一种增量式文本软聚类算法[J]. 西安交通大学学报, 2007,41(4):398-401+411.
冯中慧, 鲍军鹏, 沈钧毅. Incremental Algorithm of Text Soft Clustering[J]. 2007, 41(4): 398-401+411.
针对传统文本聚类算法时间复杂度较高
而与距离无关的算法又不适用于动态、变化的文本集等问题
提出了一种基于语义序列的增量式文本软聚类算法.该算法考虑了长文本的多主题特性
并利用语义序列相似关系计算相似语义序列集合的覆盖度
同时将每次选择的具有最小熵重叠值的候选类作为一个结果聚类
这样在整个聚类的过程中大大减小了文本向量空间的维数
缩短了计算时间.由于所提算法的语义序列只与文本自身相关
所以它适用于增量式聚类.实验结果表明
算法的聚类精度高于同条件下的其他聚类算法
尤其适合于长文本集的软聚类.
Focusing on the problems that the text clustering has high time complexity
the algorithms that are independent on the distance are unsuitable for dynamic and changing corpus
and the multi-subject characteristics of a single text cannot be considered in traditional algorithms
an incremental algorithm of text soft clustering based on semantic sequence is proposed
in which the clustering candidate with minimum entropy overlap value is selected as a result cluster by using similarity relation of semantic sequences and calculating the coverage of similarity semantic sequences set. The dimensions of text vector space are decreased dramatically in the clustering procedure
so the computing time can be reduced. Since the semantic sequence is only related to text
it is available for incremental clustering. The comparison of experimental results shows that the algorithm can achieve higher precision than other algorithms under same conditions
especially for soft clustering of long texts set.
Boley D, Gini M,Gross R,et al. Partitioning-based clustering for web document categorization [J]. Decision Support Systems, 1999, 27(3):329-341.
Zamir O, Etzioni O. Web document clustering: a feasibility demonstration [C]∥Proceedings of the 19th ACM SIGIR Conference on Research and Development in Information Retrieval. New York: ACM Press, 1998:46-54.
Dhillon I S, Guan Y,Kogan J. Co-clustering documents and words using bipartite spectral graph partitioning [C]∥Proceedings of the 7th ACM SIGKDD Conference on Knowledge Discovery and Data Mining. New York: ACM Press, 2001:269-274.
Bao J P, Shen J Y, Liu X D, et al. Semantic sequence kin: a method of document copy detection [C]∥Proceedings of the 8th Pacific-Asia Conference on Knowledge Discovery and Data Mining. Berlin:Springer-Verlag, 2004:529-538.
Beil F, Ester M, Xu X W. Frequent term-based text clustering [C]∥Proceedings of the 8th ACM SIGKDD Conference on Knowledge Discovery and Data Mining. New York: ACM Press, 2002:436-442.
Cover T M, Thomas J A. Elements of information theory [M]. New York: Wiley, 1991.
Slonim N, Tishby N. Document clustering using word clusters via the information bottleneck method [C]∥Proceedings of the 21st ACM SIGIR Conference on Research and Development in Information Retrieval. New York: ACM Press, 2000:208-215.
0
浏览量
5
下载量
2
CSCD
关联资源
相关文章
相关作者
相关机构
京公网安备11010802024621