The current feature selection algorithms based on the neighborhood rough set(NRS)model are unable to evaluate numerical dataset directly
a discretization procedure becomes necessary to transform the datasets into discrete forms
but inevitably leads to useful decision information loss. To solve this difficulty
a feature selection algorithm based on the neighborhood effective information rate is proposed. In view point of granulated neighborhood
the relation between the decision discernibility and the decision distribution is analyzed
and the neighborhood decision certainty(N
c
)is defined to indicate the degree of distinguishing capability in each individual neighborhood granule. The neighborhood decision distinguishing rate(NDDR)of the feature subset
which evaluates the ability of the subspace to approximate decision space
is est
ablished based on the sum of the N
c
values of the information granules induced by the corresponding feature space. Then the nominal and numerical datasets can be integrated into the same feature selection algorithm framework. The simulation and application illustrate that the proposed algorithm outperforms the other NRS-based ones.
关键词
Keywords
references
LIU H, YU L. Toward integrating feature selection algorithms for classification and clustering [J]. IEEE Transactions on Knowledge and Data Engineering, 2005, 17(4): 491-502.
GUYON I, ELISSEEFF A. An introduction to variable and feature selection [J]. The Journal of Machine Learning Research, 2003, 3(7/8): 1157-1182.
MITRA P, MURTHY C, PAL S. Unsupervised feature selection using feature similarity [J]. IEEE Transactions on Pattern Analysis and Machine Intelligence, 2002, 24(3): 301-312.
DASH M, LIU H. Consistency-based search in feature selection [J]. Artificial Intelligence, 2003, 151(1/2): 155-176.
DASH M, CHOI K, SCHEUERMANN P, et al. Feature selection for clustering: a filter solution [C]∥Second IEEE International Conference on Data Mining. Piscataway, NJ, USA:IEEE, 2002: 115-122.
HO T, BASU M. Complexity measures of supervised classification problems [J]. IEEE Transactions on Pattern Analysis and Machine Intelligence, 2002, 24(3): 289-300.
CHING J, WONG A, CHAN K. Class-dependent discretization for inductive learning from continuous and mixed-mode data [J]. IEEE Transactions on Pattern Analysis and Machine Intelligence, 1995, 17(7): 641-651.
POLKOWSKI L.Rough sets: mathematical foundations[M]. New York, USA: Physica-Verlag, 2002: 1-36.
HU Q, ZHANG L, CHEN D, et al. Gaussian kernel based fuzzy rough sets: model, uncertainty measures and applications [J]. International Journal of Approximate Reasoning, 2010, 51(4): 453-471.
HU Q, PEDRYCZ W, YU D, et al. Selecting discrete and continuous features based on neighborhood decision error minimization [J]. IEEE Transactions on Systems, Man, and Cybernetics: Part B Cybernetics, 2010, 40(1): 137-150.
HU Q, CHE X, ZHANG L, et al. Feature evaluation and selection based on neighborhood soft margin [J]. Neurocomputing, 2010, 73(10/11/12): 2114-2124.
HU Q, YU D, LIU J, et al. Neighborhood rough set based heterogeneous feature subset selection [J]. Information Sciences, 2008, 178(18): 3577-3594.
FRANK A, ASUNCION A. UCI machine learning repository [DB/OL]. [2010-08-22]. http:∥archive.ics.uci.edu/ml.
WANG J, WU X, ZHANG C. Support vector machines based on k-means clustering for real-time business intelligence systems [J]. International Journal of Business Intelligence and Data Mining, 2005, 1(1): 54-64.
KWAK N, CHOI C. Input feature selection for classification problems [J]. IEEE Transactions on Neural Networks, 2002, 13(1): 143-159.
QUINLAN R. Data mining tools See5 and C5.0 [EB/OL]. [2012-02-20]. http:∥www.rulequest.com/see5-info.html.
GHEYAS I, SMITH L. Feature subset selection in large dimensionality domains [J]. Pattern Recognition, 2010, 43(1): 5-13.
张晓东. 先进控制技术在选矿过程控制中的应用研究 [D]. 沈阳: 东北大学, 1999.
KARR C L, WECK B. Computer modeling of mineral processing equipment using fuzzy mathematics [J]. Minerals Engineering, 1996, 9(2): 183-194.