Focusing on the default that in the clustering algorithm based on density the objects are processed from high-density area to low-density area
thereby
the objects with low density can not be identified
a novel clustering algorithm based on boundary identification is proposed through discussing the boundary identification existed in clustering process. The main idea of the algorithm is that the objects belonging to an accumulated cluster have higher priority than the density priority
i.e. objects belonging to the same accumulated cluster will be clustered first before next clustering is processed. When the object extension reaches the boundary of cluster
the extension is stopped and turns to other direction. This method can maximize each cluster. After analyzing the density features of the cluster boundary
a boundary identification rule is created and data is clustered according to it. Experiments with synthetic data set and UCI KDD data sets demonstrate that the proposed algorithm is specially suited for processing objects with low-density
and the new algorithm can improve the performance by 4% while keeping same time complexity compared to other algorithms.
关键词
Keywords
references
Fayyad M, Piatetsky-Shapiro G, Smyth P. From data mining to knowledge discovery: an overview [C]∥Advances in Knowledge Discovery and Data Mining. Menlo Park, USA: AAAI Press, 1996:1-34.
Ester M, Kriegel H P, Sander J, et al. A density based algorithm for discovering clusters in large spatial databases with noise [C]∥Proceedings of 2nd International Conference on Knowledge Discovery and Data Mining. Oregon Portland: AAAI Press, 1996:226-231.
Guha S, Rastogi R, Shim K. CURE: an efficient clustering algorithm for large databases [C]∥Proceedings of ACM SIGMOD International Conference on Management of Data. New York: ACM Press, 1998:73-84.
Ma Shuai, Wang Tengjiao, Tang Shiwei, et al. A fast clustering algorithm based on reference and density [J]. Journal of Software, 2003, 14(6):1089-1095.
Ankerst M, Breunig M, Kriegel H P, et al. OPTICS: ordering points to identify the clustering structure [C]∥Proceedings of ACM SIGMOD International Conference on Management of Data. New York: ACM Press, 1999:49-60.
Sun Xuegang, Chen Qunfang, Ma Liang. Study on topic-based Web clustering [J]. Journal of Chinese Information Processing, 2003, 17(3):21-26.
Ayad H, Kamel M. Topic discovery from text using aggregation of different clustering methods [C]∥Proceedings of the 15th Conference of the Canadian Society for Computational Studies of Intelligence on Advances in Artificial Intelligence. Heidelberg, Germany: Springer-Verlag, 2002:161-175.