北京邮电大学可信分布式计算与服务教育部重点实验室,北京,100876
网络首发:2018-10-10,
纸质出版:2018
移动端阅览
张俐, 王枞, 郭文明. 利用近似马尔科夫毯的最大相关最小冗余特征选择算法[J]. 西安交通大学学报, 2018,52(10):141-145.
A Feature Selection Algorithm for Maximum Relevance Minimum Redundancy Using Approximate Markov Blanket[J]. 2018, 52(10): 141-145.
张俐, 王枞, 郭文明. 利用近似马尔科夫毯的最大相关最小冗余特征选择算法[J]. 西安交通大学学报, 2018,52(10):141-145. DOI: 10.7652/xjtuxb201810019.
A Feature Selection Algorithm for Maximum Relevance Minimum Redundancy Using Approximate Markov Blanket[J]. 2018, 52(10): 141-145. DOI: 10.7652/xjtuxb201810019.
针对高维数据集中冗余特征或无关特征降低机器学习模型分类准确率的问题
提出了一种基于近似马尔科夫毯的特征选择(nmRMR)算法。该算法首先利用最大相关最小冗余的准则进行特征相关性排序; 采用近似马尔科夫毯算法对冗余特征或者无关特征进行删除
并最大程度地提高特征之间的相关性从而获得最优特征子集。在UCI的8个公开数据集上对比的实验结果表明:与mRMR算法相比
本文算法所选择出的特征子集数平均减少了6.875个
平均分类准确率提高了0.78%; 与FullSet算法相比
本文算法所选择出的特征子集数平均减少了20.56个
平均分类准确率提高了1.88%; 与FCBF算法相比
本文算法所选择出的特征子集数平均减少了3.187 5个
平均分类准确率提高了0.825%; 本文算法总体优于其他算法。
To solve the problem that redundancy or irrelevant features in high-dimensional datasets reduce the classification accuracy of machine learning model
a feature selection algorithm based on approximate Markov blanket is proposed and named as normal max-relevance and min-redundancy(nmRMR)algorithm. Firstly
the algorithm uses the criteria of maximum relevance and minimum redundancy to perform feature relevance ranking. Then
it adopts the approximate Markov blanket to remove redundant features or irrelevant features
and maximize the correlation between features to obtain the optimal feature subset. Experimental results on UCI's eight open datasets show that: the proposed nmRMR algorithm achieves on average 6.875
20.56 and 3.187 5 reduction in the selected number of feature subsets
as well as 0.78%
1.88% and 0.825% improvement in the average classification accuracy
compared with the mRMR algorithm
the FullSet algorithm
and the FCBF algorithm
respectively. It is concluded that the proposed nmRMR algorithm is superior to other algorithms.
BENNASAR M, HICKS Y, SETCHI R. Feature selection using joint mutual information maximisation [J]. Expert Systems with Applications, 2015, 42(22): 8520-8532.
ZHANG Y, ZHANG Z. Feature subset selection with cumulate conditional mutual information minimization [J]. Expert Systems with Applications, 2012, 39(5): 6078-6088.
LIU Chuan, WANG Wenyong, ZHAO Qiang, et al. A new feature selection method based on a validity index of feature subset [J]. Pattern Recognition Letters, 2017, 92: 1-8.
SUN Xin, LIU Yanheng, LI Jin, et al. Feature evaluation and selection with cooperative game theory [J]. Pattern Recognition, 2012, 45(8): 2992-3002.
PENG H, LONG F, DING C. Feature selection based on mutual information criteria of max-dependency, max-relevance, and min-redundancy [J]. IEEE Transactions Pattern Analysis and Machine Intelligence, 2005, 27(8): 1226-1238.
ZHOU P, HU X, LI P, et al. Online feature selection for high-dimensional class-imbalanced data [J]. Knowledge-Based Systems, 2017, 136(15): 187-199.
GUO Q, ZHANG M. Implement web learning environment based on data mining [J]. Knowledge-Based Systems, 2009, 22(6): 439-442.
WANG Y, WANG J, LIAO H, et al. An efficient semi-supervised representatives feature selection algorithm based on information theory [J]. Pattern Recognition, 2017, 61(1): 511-523.
KWAK N, CHOI C H. Input feature selection for classification problems [J]. IEEE Transactions on Neural Networks, 2002, 13(1): 143-159.
KOLLER D, SAHAMI M. Toward optimal feature selection [C]∥Proceedings of the 13th International Conference on International Conference on Machine Learning. San Francisco, CA, USA: Morgan Kaufmann Publishers Inc., 1996: 284-292.
崔自峰, 徐宝文, 张卫丰, 等. 一种近似Markov Blanket最优特征选择算法 [J]. 计算机学报, 2007, 30(12): 2074-2081.
CUI Zifeng, XU Baowen, ZHANG Weifeng, et al. An approximate Markov blanket feature selection algorithm [J]. Chinese Journal of Computers, 2007, 30(12): 2074-2081.
KURGAN L A, CIOS K J. CAIM discretization algorithm [J]. IEEE Transactions on Knowledge Data Engineering, 2004, 16(2): 145-153.
RISH I. An empirical study of the naive Bayes classifier [J]. Journal of Universal Computer Science, 2001, 1(2): 127: 41-46.
BERMEJO P, OSSA L D L, GAMEZ J A, et al. Fast wrapper feature subset selection in high-dimensional datasets by means of filter re-ranking [J]. Knowledge-Based Systems, 2012, 25(1): 35-44.
GOMEZ-VERDEJO V, VERLEYSEN M, FLEURY J. Information-theoretic feature selection for functional data classification [J]. Neurocomputing, 2009, 72(16/17/18): 3580-3589.
杨宏晖,王芸,孙进才,等.融合样本选择与特征选择的AdaBoost支持向量机集成算法.2014,48(12):63-68.[doi:10.7652/xjtuxb201412010]
诸文智,司刚全,张彦斌.采用邻域决策分辨率的特征选择算法.2013,47(2):20-27.[doi:10.7652/xjtuxb201302004]
栗茂林,梁霖,王孙安,等.结合交叠区异点统计和相关分析的免疫克隆特征选择方法.2012,46(5):50-56.[doi:10.7652/xjtuxb201205009]
豆增发,高琳.利用膜粒子群优化和信息熵的医学文本特征选择.2012,46(4):45-51.[doi:10.7652/xjtuxb201204008]
杨宏晖,戴健,孙进才,等.用于水声目标识别的自适应免疫特征选择算法.2011,45(12):28-32.[doi:10.7652/xjtuxb 201112006]
薛峰,周亚东,高峰,等.一种突发性热点话题在线发现与跟踪方法.2011,45(12):64-69.[doi:10.7652/xjtuxb201112 012]
马超,陈西宏,徐宇亮,等.广义邻域粗集下的集成特征选择及其选择性集成算法.2011,45(6):34-39.[doi:10.7652/xjtuxb201106006]
裴晓梅,郑崇勋.基于Fisher判据时频分析的运动相关脑电特征选择及优化.2008,42(8):1026-1030.[doi:10.7652/xjtuxb200808021]
朱虎明,焦李成.基于免疫记忆克隆的特征选择.2008,42(6):679-682.[doi:10.7652/xjtuxb200806007]
崔舒宁,朱丹军,冯博琴,等.结合受控词汇表的生物基因本体标注与分类.2008,42(2):171-174.[doi:10.7652/xjtuxb 200802010]
0
浏览量
5
下载量
8
CSCD
关联资源
相关文章
相关作者
相关机构
京公网安备11010802024621