西安交通大学计算机科学与技术系,西安,710049
网络首发:2007-12-10,
纸质出版:2007
移动端阅览
张选平, 祝兴昌, 马琮. 一种基于边界识别的聚类算法[J]. 西安交通大学学报, 2007,41(12):1387-1390+1395.
张选平, 祝兴昌, 马琮. Clustering Algorithm Based on Boundary Identification[J]. 2007, 41(12): 1387-1390+1395.
针对基于密度的聚类算法由高密度区到低密度区的处理顺序所带来的不能识别低密度对象类别的缺陷
通过对聚类过程中可能存在的边界识别进行讨论
提出了一种基于边界识别的聚类算法.该算法的思想是:同簇优先权高于密度优先权
即在选择下一个对象进行聚类时
在已聚类的对象中优先选择同一簇的对象
当对象沿某一方向扩展到达簇边界时停止扩展
转而向其他方向扩展
这种处理顺序能使得类别最大化.通过分析簇边界的密度变化特征
建立了边界识别准则
并根据该准则对数据进行聚类.通过在合成数据和美国加州大学提供的知识挖掘数据库数据集上的实验结果表明
所提算法能有效地处理低密度区域的数据
与识别聚类结构的对象排序算法相比
聚类效果可提高4%左右
而时间性能相当.
Focusing on the default that in the clustering algorithm based on density the objects are processed from high-density area to low-density area
thereby
the objects with low density can not be identified
a novel clustering algorithm based on boundary identification is proposed through discussing the boundary identification existed in clustering process. The main idea of the algorithm is that the objects belonging to an accumulated cluster have higher priority than the density priority
i.e. objects belonging to the same accumulated cluster will be clustered first before next clustering is processed. When the object extension reaches the boundary of cluster
the extension is stopped and turns to other direction. This method can maximize each cluster. After analyzing the density features of the cluster boundary
a boundary identification rule is created and data is clustered according to it. Experiments with synthetic data set and UCI KDD data sets demonstrate that the proposed algorithm is specially suited for processing objects with low-density
and the new algorithm can improve the performance by 4% while keeping same time complexity compared to other algorithms.
Fayyad M, Piatetsky-Shapiro G, Smyth P. From data mining to knowledge discovery: an overview [C]∥Advances in Knowledge Discovery and Data Mining. Menlo Park, USA: AAAI Press, 1996:1-34.
Ester M, Kriegel H P, Sander J, et al. A density based algorithm for discovering clusters in large spatial databases with noise [C]∥Proceedings of 2nd International Conference on Knowledge Discovery and Data Mining. Oregon Portland: AAAI Press, 1996:226-231.
Guha S, Rastogi R, Shim K. CURE: an efficient clustering algorithm for large databases [C]∥Proceedings of ACM SIGMOD International Conference on Management of Data. New York: ACM Press, 1998:73-84.
马帅,王腾蛟,唐世渭,等.一种基于参考点和密度的快速聚类算法 [J].软件学报,2003,14(6):1089-1095.
Ma Shuai, Wang Tengjiao, Tang Shiwei, et al. A fast clustering algorithm based on reference and density [J]. Journal of Software, 2003, 14(6):1089-1095.
Ankerst M, Breunig M, Kriegel H P, et al. OPTICS: ordering points to identify the clustering structure [C]∥Proceedings of ACM SIGMOD International Conference on Management of Data. New York: ACM Press, 1999:49-60.
孙学刚, 陈群芳, 马亮. 基于主题的Web文档聚类研究 [J]. 中文信息学报,2003,17(3):21-26.
Sun Xuegang, Chen Qunfang, Ma Liang. Study on topic-based Web clustering [J]. Journal of Chinese Information Processing, 2003, 17(3):21-26.
Ayad H, Kamel M. Topic discovery from text using aggregation of different clustering methods [C]∥Proceedings of the 15th Conference of the Canadian Society for Computational Studies of Intelligence on Advances in Artificial Intelligence. Heidelberg, Germany: Springer-Verlag, 2002:161-175.
0
浏览量
6
下载量
0
CSCD
关联资源
相关文章
相关作者
相关机构
京公网安备11010802024621