西安交通大学电气工程学院,西安,710049
网络首发:2013-02-10,
纸质出版:2013
移动端阅览
诸文智, 司刚全, 张彦斌. 采用邻域决策分辨率的特征选择算法[J]. 西安交通大学学报, 2013,47(2):20-27.
Feature Selection Algorithm Based on Neighborhood Decision Distinguishing Rate[J]. 2013, 47(2): 20-27.
诸文智, 司刚全, 张彦斌. 采用邻域决策分辨率的特征选择算法[J]. 西安交通大学学报, 2013,47(2):20-27. DOI: 10.7652/xjtuxb201302004.
Feature Selection Algorithm Based on Neighborhood Decision Distinguishing Rate[J]. 2013, 47(2): 20-27. DOI: 10.7652/xjtuxb201302004.
针对目前基于粗糙集模型的特征选择算法无法直接应用于数值型数据、必须经过离散化过程而造成决策信息丢失的问题
提出了一种基于邻域决策分辨率的特征选择算法。该算法根据邻域信息粒中决策分布与其分类能力间的关系
提出了邻域决策确定性(N
c
)来衡量单个信息粒的决策分辨能力; 并根据特征向量空间上所有信息粒所具有的N
c
累加值
定义了邻域决策分辨率作为特征子集上决策可分辨性的量度
从而将名义型和数值型数据统一在同一特征选择算法框架下。仿真实验和实际应用的结果表明
该算法性能优于目前主流基于邻域粗糙集的特征选择方法。
The current feature selection algorithms based on the neighborhood rough set(NRS)model are unable to evaluate numerical dataset directly
a discretization procedure becomes necessary to transform the datasets into discrete forms
but inevitably leads to useful decision information loss. To solve this difficulty
a feature selection algorithm based on the neighborhood effective information rate is proposed. In view point of granulated neighborhood
the relation between the decision discernibility and the decision distribution is analyzed
and the neighborhood decision certainty(N
c
)is defined to indicate the degree of distinguishing capability in each individual neighborhood granule. The neighborhood decision distinguishing rate(NDDR)of the feature subset
which evaluates the ability of the subspace to approximate decision space
is est
ablished based on the sum of the N
c
values of the information granules induced by the corresponding feature space. Then the nominal and numerical datasets can be integrated into the same feature selection algorithm framework. The simulation and application illustrate that the proposed algorithm outperforms the other NRS-based ones.
LIU H, YU L. Toward integrating feature selection algorithms for classification and clustering [J]. IEEE Transactions on Knowledge and Data Engineering, 2005, 17(4): 491-502.
GUYON I, ELISSEEFF A. An introduction to variable and feature selection [J]. The Journal of Machine Learning Research, 2003, 3(7/8): 1157-1182.
MITRA P, MURTHY C, PAL S. Unsupervised feature selection using feature similarity [J]. IEEE Transactions on Pattern Analysis and Machine Intelligence, 2002, 24(3): 301-312.
DASH M, LIU H. Consistency-based search in feature selection [J]. Artificial Intelligence, 2003, 151(1/2): 155-176.
DASH M, CHOI K, SCHEUERMANN P, et al. Feature selection for clustering: a filter solution [C]∥Second IEEE International Conference on Data Mining. Piscataway, NJ, USA:IEEE, 2002: 115-122.
HO T, BASU M. Complexity measures of supervised classification problems [J]. IEEE Transactions on Pattern Analysis and Machine Intelligence, 2002, 24(3): 289-300.
CHING J, WONG A, CHAN K. Class-dependent discretization for inductive learning from continuous and mixed-mode data [J]. IEEE Transactions on Pattern Analysis and Machine Intelligence, 1995, 17(7): 641-651.
POLKOWSKI L.Rough sets: mathematical foundations[M]. New York, USA: Physica-Verlag, 2002: 1-36.
HU Q, ZHANG L, CHEN D, et al. Gaussian kernel based fuzzy rough sets: model, uncertainty measures and applications [J]. International Journal of Approximate Reasoning, 2010, 51(4): 453-471.
HU Q, PEDRYCZ W, YU D, et al. Selecting discrete and continuous features based on neighborhood decision error minimization [J]. IEEE Transactions on Systems, Man, and Cybernetics: Part B Cybernetics, 2010, 40(1): 137-150.
HU Q, CHE X, ZHANG L, et al. Feature evaluation and selection based on neighborhood soft margin [J]. Neurocomputing, 2010, 73(10/11/12): 2114-2124.
HU Q, YU D, LIU J, et al. Neighborhood rough set based heterogeneous feature subset selection [J]. Information Sciences, 2008, 178(18): 3577-3594.
FRANK A, ASUNCION A. UCI machine learning repository [DB/OL]. [2010-08-22]. http:∥archive.ics.uci.edu/ml.
WANG J, WU X, ZHANG C. Support vector machines based on k-means clustering for real-time business intelligence systems [J]. International Journal of Business Intelligence and Data Mining, 2005, 1(1): 54-64.
KWAK N, CHOI C. Input feature selection for classification problems [J]. IEEE Transactions on Neural Networks, 2002, 13(1): 143-159.
QUINLAN R. Data mining tools See5 and C5.0 [EB/OL]. [2012-02-20]. http:∥www.rulequest.com/see5-info.html.
GHEYAS I, SMITH L. Feature subset selection in large dimensionality domains [J]. Pattern Recognition, 2010, 43(1): 5-13.
张晓东. 先进控制技术在选矿过程控制中的应用研究 [D]. 沈阳: 东北大学, 1999.
KARR C L, WECK B. Computer modeling of mineral processing equipment using fuzzy mathematics [J]. Minerals Engineering, 1996, 9(2): 183-194.
0
浏览量
4
下载量
2
CSCD
关联资源
相关文章
相关作者
相关机构
京公网安备11010802024621