西安交通大学机械制造系统工程国家重点实验室,西安,710049
网络首发:2017-12-10,
纸质出版:2017
移动端阅览
姜洪权 1, 王岗 2, 高建民 1, 等. 一种适用于高维非线性特征数据的聚类算法及应用[J]. 西安交通大学学报, 2017,51(12):49-55+90.
A Clustering Algorithm for High-Dimensional Nonlinear Feature Data with Applications[J]. 2017, 51(12): 49-55+90.
姜洪权 1, 王岗 2, 高建民 1, 等. 一种适用于高维非线性特征数据的聚类算法及应用[J]. 西安交通大学学报, 2017,51(12):49-55+90. DOI: 10.7652/xjtuxb201712008.
A Clustering Algorithm for High-Dimensional Nonlinear Feature Data with Applications[J]. 2017, 51(12): 49-55+90. DOI: 10.7652/xjtuxb201712008.
针对高维数据聚类分析中数据之间具有多种非线性特征关系
导致数据分布不均、传统相似性度量失效及结果类中心难以精准表征等问题
提出了一种基于核主元分析(KPCA)与密度聚类(DBSCAN)的高维非线性特征数据聚类分析技术。首先
为有效提取高维数据的非线性特征
利用KPCA理论将原始数据映射到更高维数据空间
利用主元分析获得数据变化的方向集合
并进行降维分析; 然后
通过重新定义数据样本在主元空间的相似性距离对传统DBSCAN聚类方法进行改进
并利用3δ统计理论对各簇中心的进行表征
从而实现高维数据的精确分类与类中心知识表达。以实际高血压患者群体聚类问题为例对方法进行了有效性验证
实验表明
所提方法可以有效获取原始数据的非线性特征
实现患者个体特征群体的有效划分及簇类中心知识的表达
解决传统DBSCAN聚类方法对高维数据不适用的问题。
Aiming at the problems caused by the nonlinear relations between the attributes of high dimensional data in cluster analysis
such as uneven distribution of data
invalidation of traditional similarity measures and difficulty of accurate representation of the result class
a clustering algorithm for high dimensional nonlinear feature data is proposed based on kernel principal component analysis(KPCA)and density clustering(DBSCAN). To extract the nonlinear characteristics of high dimensional data
the KPCA theory is adopted to map the original to a higher dimensional data space
thus a set of directions in principal component space-PCS for extracting the nonlinear characteristics of data and reduced dimensions can be obtained. The similarity distance of data in PCS is defined to improve the traditional DBSCAN clustering algorithm and 3δ statistical theory is used to characterize the clustering results. A case of hypertension group clustering is provided to illustrate the feasibility of the proposed method
and the results show that the proposed method can effectively obtain the nonlinear characteristics of the high dimensional data and realize cluster analysis and cluster center knowledge expression to solve the difficulties in the traditional DBSCAN clustering method for cluster analysis of high dimensional data.
王骏, 王士同, 邓赵红. 聚类分析研究中的若干问题 [J]. 控制与决策, 2012, 27(3): 321-327.
WANG Jun, WANG Shitong, DENG Zhaohong. Survey on challenges in clustering analysis research [J]. Control and Decision, 2012, 27(3): 321-327.
刘红岩, 陈剑, 陈国青. 数据挖掘中的数据分类算法综述 [J]. 清华大学学报(自然科学版), 2002, 42(6): 727-730.
LIU Hongyan, CHEN Jian, CHEN Guoqing. Review of classification algorithms for data mining [J]. Journal of Tsinghua University(Science and Technology), 2002, 42(6): 727-730.
蔡颖琨, 谢昆青, 马修军. 屏蔽了输入参数敏感性的DBSCAN改进算法 [J]. 北京大学学报(自然科学版), 2014, 40(3): 480-486.
CAI Yingkun, XIE Kunqing, MA Xiujun. An improved DBSCAN algorithm which is insensitive to input parameters [J]. Acta Scientiarum Naturalium Universitatis Pekinensis, 2014, 40(3): 480-486.
和亚丽. 基于高维空间的聚类技术研究 [D]. 太原: 中北大学, 2005: 8-13.
ESTER M, KRIEGEL H P, SANDER J, et al. A density-based algorithm for discovering clusters in large spatial databases with noise [C]∥Proc 2nd Int Conf on Knowledge Discovery and Data Mining(KDD-96). Portland, USA: ACM Press, 1996: 226-231.
于亚飞, 周爱武. 一种改进的DBSCAN密度算法 [J]. 计算机技术与发展, 2011, 21(2): 30-33.
YU Yafei, ZHOU Aiwu. An improved algorithm of DBSCAN [J]. Computer Technology and Development, 2011, 21(2): 30-33.
贺玲, 吴玲达. 高维空间中数据的相似性度量 [J]. 数学的实践与认识, 2006, 36(9): 189-193.
HE Ling, WU Lingda. Similarity measurement of data in high-dimensional spaces [J]. Mathematics in Practice and Theory, 2006, 36(9): 189-193.
FRIEDMAN J H. Flexible metric nearest neighbor classification [EB/OL]. [2017-04-15]. http: ∥docs. salford-systems.com/flexmet.pdf.
王瀛, 郭雷, 梁楠. 基于优选样本的KPCA高光谱图像降维方法 [J]. 光子学报, 2011, 40(6): 847-851.
WANG Ying, GUO Lei, LIANG Nan. A dimensionality reduction method based on KPCA with optimized sample set for hyperspectral image [J]. Acta Photonica Sinica, 2011, 40(6): 847-851.
杨喆祾. 基于经验模态分解的城市供水水质异常事件检测方法研究 [D]. 杭州: 浙江大学, 2016: 39-40.
高智勇, 梁银林. 基于集成熵KPCA的复杂机电系统状态监测方法 [J]. 计算机集成制造系统, 2015, 21(5): 1327-1333.
GAO Zhiyong, LIANG Yinlin. State monitoring of complex electromechanical system based on integrated entropy of KPCA [J]. Computer Integrated Manufacturing Systems, 2015, 21(5): 1327-1333.
杨风召, 朱扬勇. 一种有效的量化交易数据相似性搜索方法 [J]. 计算机研究与发展, 2004, 41(2): 361-368.
YANG Fengzhao, ZHU Yangyong. An efficient method for similarity search on quantitative transaction data [J]. Journal of Computer Research and Development, 2004, 41(2): 361-368.
杨燕, 靳蕃. 聚类有效性评价综述 [J]. 计算机应用研究, 2008, 25(6): 1630-1632.
YANG Yan, JIN Fan. Survey of clustering validity evaluation [J]. Application Research of Computers, 2008, 25(6): 1630-1632.
0
浏览量
5
下载量
2
CSCD
关联资源
相关文章
相关作者
相关机构
京公网安备11010802024621