1. 西安交通大学电子与信息工程学院,西安,710049
2. 新疆大学信息科学与工程学院,乌鲁木齐,830046
网络首发:2011-08-10,
纸质出版:2011
移动端阅览
王羡慧 1, 3, 覃征 1, 等. 采用仿射传播的聚类集成算法[J]. 西安交通大学学报, 2011,45(8):1-6.
Cluster Ensemble Algorithm Using Affinity Propagation[J]. 2011, 45(8): 1-6.
针对K均值聚类随机初始聚类中心导致的聚类结果不稳定问题
提出一种基于仿射传播的聚类集成算法.该算法把每个聚类集成的成员个体结果看成是原始数据的一个属性
然后在其基础上对聚类成员个体的聚类结果进行加权集成
集成算法采用简单高效的仿射传播聚类
并且提出了直接集成、利用平均规范化互信息(NMI)和聚类有效性Silhouette指标进行加权集成.最后
运用Hungarian算法对仿射传播聚类集成的结果进行类别标签的统一和匹配.在加州大学尔湾分校数据集上进行了实验
结果表明
与集成前的K均值聚类及其他聚类集成算法相比
该算法能有效地提高聚类结果的准确性、鲁棒性和稳定性
建立起来的聚类集成算法具有良好的扩展性和灵活性
而且简单有效.
The result of K-means cluster is instable for random initial clustering centers. A cluster ensemble algorithm based on affinity propagation is proposed
where the result of each cluster individual is regarded as a property of the original data. Following the new properties sets
the results of each cluster individual are carried out to a weighted ensemble
and simple and efficient affinity propagation cluster is chosen in the ensemble algorithm. Furthermore the direct ensemble
the ensemble to weighted ensemble from average normalized mutual information(NMI)and cluster validation indexes Silhouette are uniformly proposed. Finally
Hungarian algorithm is employed to unify and match the category labels for the results of affinity propagation cluster. The results of experiments on University of California Irvine data sets show the higher efficiency for improving the accuracy
robustness and stability of cluster results than the K means clustering before combination and the other clustering ensemble algorithms. The clustering ensemble algorithm gets more extendable and flexible.
XU R, WUNSCH D. Survey of clustering algorithms[J]. IEEE Transactions on Neural Networks, 2005, 16(3): 645-678.
OMRAN M G H, ENGELBRECHT A P, SALMAN A. An overview of clustering methods[J]. Intelligent Data Analysis, 2007, 11(6): 583-605.
孙吉贵,刘杰,赵连宇. 聚类算法研究[J]. 软件学报, 2008, 19(1):48-61.
SUN Jigui, LIU Jie, ZHAO Lianyu. Clustering algorithms research[J]. Journal of Software, 2008, 19(1):48-61.
MACQUEEN J. Some methods for classification and analysis of multivariate observations[C]∥Proceedings of the Fifth Berkeley Symposium on Mathematical Statistics and Probability. Berkeley, California,USA: University of California Press, 1967: 281-297.
徐森, 卢志茂, 顾国昌. 解决文本聚类集成问题的两个谱算法[J]. 自动化学报, 2009, 35(7):997-1002
XU Sen, LU Zhimao, GU Guochang. Two spectral algorithms for ensembling document clusters[J]. Acta Automatica Sinica, 2009, 35(7): 997-1002.
FRED A, JAIN A. Combining multiple clusterings using evidence accumulation[J]. IEEE Transactions on Pattern Analysis and Machine Intelligence, 2005, 27(6): 835-850.
ZHOU Z H, TANG W. Clusterer ensemble[J]. Knowledge-Based Systems, 2006, 19(1): 77-83.
罗会兰, 孔繁胜, 李一啸. 聚类集成中的差异性度量研究[J]. 计算机学报, 2007, 30(8): 1315-1324.
LUO Huilan, KONG Fansheng, LI Yixiao.An analysis of diversity measures in clustering ensembles[J]. Chinese Journal of Computers, 2007, 30(8): 1315-1324.
STREHL A, GHOSH J. Cluster ensembles: a knowledge reuse framework for combining multiple partitions[J]. The Journal of Machine Learning Research, 2002(3): 583-617.
王红军, 李志蜀,成飚, 等. 基于隐含变量的聚类集成模型[J]. 软件学报, 2009, 20(4): 825-833.
WANG Hongjun,LI Zhishu,CHENG Biao, et al. A latent variable mode for cluster ensemble [J]. Journal of Software, 2009, 20(4): 825-833.
FREY B J, DUECK D. Clustering by passing messages between data points[J]. Science, 2007, 315(5814): 972-976.
FREY B J, DUECK D. Response to comment on “clustering by passing messages between data points”[J]. Science, 2008, 319(5864): 2.
MEZARD M. Computer science: where are the exemplars?[J]. Science, 2007, 315(5814): 949-951.
FRED A. Finding consistent clusters in data partitions[M]. Heidelberg,German: Springer, 2001:309-318.
HRUSCHKA E R, CAMPELLO R, FREITAS A A, et al. A survey of evolutionary algorithms for clustering[J]. IEEE Transactions on Systems, Man and Cybernetics:Part C Applications and Reviews, 2009, 39(2): 133-155.
KAUFMAN L, ROUSSEEUW P J. Finding groups in data: an introduction to cluster analysis[M]. New York: Wiley, 1990:7-64.
KUHN H W. The Hungarian method for the assignment problem[J]. Naval Research Logistics Quarterly, 1955, 2(2): 83-97.
MODHA D S, SPANGLER W S. Feature weighting in k-means clustering[J]. Machine Learning, 2003, 52(3): 217-237.
FRANK A,ASUNCION A. UCI machine learning repository [EB/OL]. [2010-12-22]. http:∥archive.ics.uci.edu/ml.
0
浏览量
4
下载量
6
CSCD
关联资源
相关文章
相关作者
相关机构
京公网安备11010802024621