Conventional association rule-based categorization methods have bottleneck in improving classifier's accuracy
since these methods only consider the rule confidence degree and use the pruning technique. A novel method to solve this problem is proposed
and is called associative rule-based classifier aggregating with category similarity(AACS). The method adopts the modified chi-square statistical technique to extract feature terms from each category
and employs the CR-tree to store classification rules. Algorithms to construct and to match CR-tree are proposed. Inner-product is used to calculate the similarity between the category sub vector of the text and the category feature vector
and then is aggregated with the rules' confidence degree to serve as the foundation of text categorization. Experimental results show that the method presented achieves a micro-average value of categorization 92.42% with extracting only 30 feature terms
which is better than the results of AWOPR
KNN
and Bayes classifiers. And the time complexity of the method is the same as that of AWOPR
indicating that the cost to calculate both the similarity and the aggregation is acceptable.
关键词
Keywords
references
LIU Bing, HSU W, MA Yiming. Integrating classification and association rule mining [C]∥Proceedings of the ACM International Conference on Knowledge Discovery and Data Mining. New York,USA: ACM, 1998: 80-86.
ZAÏANE O R, ANTONIE M L. Classifying text documents by associating terms with text categories [C]∥Proceedings of the 13th Australasian Database Conference. New York,USA: ACM, 2002: 215-222.
LI Wenmin, HAN Jiawei, PEI Jian. CMAR: Accurate and efficient classification based on multiple classification rules [C]∥Proceedings of the 2001 IEEE International Conference on Data Mining. Piscataway,NJ,USA: IEEE, 2001: 369-376.
CHEN Xiaoyun, CHEN Yi, WANG Lei, et al. Text categorization based on classification rules tree by frequent patterns [J]. Journal of Software, 2006, 17(5): 1017-1025.
CHEN Xiaoyun, HU Yunfa. Text association categorization based on self-adaptive weighting [J]. Journal of Chinese Computer Systems, 2007, 28(1):116-121.
SHANG Bingzhang, BAI Qingyuan. Improved association text classification based on feature weight [J]. Journal of Computer Research and Development, 2008, 45(S0): 252-256.
CAI Jinfeng, BAI Qingyuan. Association text classification of mining ItemSet significance [J]. Journal of Nanjing University, 2011, 47(5): 544-550.
BARALIS E, GARZA P. I-prune: item selection for associative classification [J]. International Journal of Intelligent Systems, 2012, 27(3): 279-299.
YANG Yiming, PEDERSON J O. A comparative study on feature selection in text categorization [C]∥Proceedings of the 14th International Conference on Machine Learning. San Francisco, CA, USA: Morgan Kaufmann, 1997: 412-420.
AGRAWAL R, SRIKANT R. Fast algorithms for mining association rules [C]∥Proceedings of the 20th VLDB Conference. San Francisco, CA, USA: Morgan Kaufmann, 1994: 487-499.
HAN Jiawei, PEI Jian, YIN Yiwen, et al. Mining frequent patterns without candidate generation: a frequent-pattern tree approach [J]. Data Mining and Knowledge Discovery, 2004, 8(1): 53-87.
SEBASTIANI F. Machine learning in automated text categorization [J]. ACM Computing Surveys, 2002, 34(1): 1-47.