The result of K-means cluster is instable for random initial clustering centers. A cluster ensemble algorithm based on affinity propagation is proposed
where the result of each cluster individual is regarded as a property of the original data. Following the new properties sets
the results of each cluster individual are carried out to a weighted ensemble
and simple and efficient affinity propagation cluster is chosen in the ensemble algorithm. Furthermore the direct ensemble
the ensemble to weighted ensemble from average normalized mutual information(NMI)and cluster validation indexes Silhouette are uniformly proposed. Finally
Hungarian algorithm is employed to unify and match the category labels for the results of affinity propagation cluster. The results of experiments on University of California Irvine data sets show the higher efficiency for improving the accuracy
robustness and stability of cluster results than the K means clustering before combination and the other clustering ensemble algorithms. The clustering ensemble algorithm gets more extendable and flexible.
关键词
Keywords
references
XU R, WUNSCH D. Survey of clustering algorithms[J]. IEEE Transactions on Neural Networks, 2005, 16(3): 645-678.
OMRAN M G H, ENGELBRECHT A P, SALMAN A. An overview of clustering methods[J]. Intelligent Data Analysis, 2007, 11(6): 583-605.
孙吉贵,刘杰,赵连宇. 聚类算法研究[J]. 软件学报, 2008, 19(1):48-61.
SUN Jigui, LIU Jie, ZHAO Lianyu. Clustering algorithms research[J]. Journal of Software, 2008, 19(1):48-61.
MACQUEEN J. Some methods for classification and analysis of multivariate observations[C]∥Proceedings of the Fifth Berkeley Symposium on Mathematical Statistics and Probability. Berkeley, California,USA: University of California Press, 1967: 281-297.
XU Sen, LU Zhimao, GU Guochang. Two spectral algorithms for ensembling document clusters[J]. Acta Automatica Sinica, 2009, 35(7): 997-1002.
FRED A, JAIN A. Combining multiple clusterings using evidence accumulation[J]. IEEE Transactions on Pattern Analysis and Machine Intelligence, 2005, 27(6): 835-850.
ZHOU Z H, TANG W. Clusterer ensemble[J]. Knowledge-Based Systems, 2006, 19(1): 77-83.
LUO Huilan, KONG Fansheng, LI Yixiao.An analysis of diversity measures in clustering ensembles[J]. Chinese Journal of Computers, 2007, 30(8): 1315-1324.
STREHL A, GHOSH J. Cluster ensembles: a knowledge reuse framework for combining multiple partitions[J]. The Journal of Machine Learning Research, 2002(3): 583-617.
WANG Hongjun,LI Zhishu,CHENG Biao, et al. A latent variable mode for cluster ensemble [J]. Journal of Software, 2009, 20(4): 825-833.
FREY B J, DUECK D. Clustering by passing messages between data points[J]. Science, 2007, 315(5814): 972-976.
FREY B J, DUECK D. Response to comment on “clustering by passing messages between data points”[J]. Science, 2008, 319(5864): 2.
MEZARD M. Computer science: where are the exemplars?[J]. Science, 2007, 315(5814): 949-951.
FRED A. Finding consistent clusters in data partitions[M]. Heidelberg,German: Springer, 2001:309-318.
HRUSCHKA E R, CAMPELLO R, FREITAS A A, et al. A survey of evolutionary algorithms for clustering[J]. IEEE Transactions on Systems, Man and Cybernetics:Part C Applications and Reviews, 2009, 39(2): 133-155.
KAUFMAN L, ROUSSEEUW P J. Finding groups in data: an introduction to cluster analysis[M]. New York: Wiley, 1990:7-64.
KUHN H W. The Hungarian method for the assignment problem[J]. Naval Research Logistics Quarterly, 1955, 2(2): 83-97.
MODHA D S, SPANGLER W S. Feature weighting in k-means clustering[J]. Machine Learning, 2003, 52(3): 217-237.
FRANK A,ASUNCION A. UCI machine learning repository [EB/OL]. [2010-12-22]. http:∥archive.ics.uci.edu/ml.