1. 西安交通大学电子与信息工程学院,西安,710049
2. 西安交通大学陕西省计算机网络重点实验室,西安,710049
网络首发:2012-12-10,
纸质出版:2012
移动端阅览
田丰 1, 桂小林 1, 杨攀 1, 等. 采用类别相似度聚合的关联文本分类方法[J]. 西安交通大学学报, 2012,46(12):6-11+122.
Associative Rule-Based Text Categorization Method Using Category Similarity[J]. 2012, 46(12): 6-11+122.
针对基于关联规则的分类方法在分类时仅考虑规则的置信度并使用规则修剪技术
导致分类器的分类精度难以进一步提高的问题
提出了一种基于类别相似度聚合的关联文本分类方法.该方法采用修改的χ
2
统计技术提取各类别的特征词; 为保证规则匹配的精度和速度
使用CR-tree存储分类规则
并给出了CR-tree的构建与匹配算法; 采用向量内积来计算文本类别分量与类别标志向量的相似度
进而使用规则置信度和类别相似度的聚合值作为文本分类的依据.基于实际网络文本的实验表明
该方法仅需提取30个特征词
分类结果的微平均值即可达到92.42%
优于未经剪枝的ARC-BC分类器及KNN、Bayes分类器; 在分类耗时方面
该方法与未经剪枝的ARC-BC分类器持平
表明该方法引入的相似度与聚合值的计算开销在可接受的范围内.
Conventional association rule-based categorization methods have bottleneck in improving classifier's accuracy
since these methods only consider the rule confidence degree and use the pruning technique. A novel method to solve this problem is proposed
and is called associative rule-based classifier aggregating with category similarity(AACS). The method adopts the modified chi-square statistical technique to extract feature terms from each category
and employs the CR-tree to store classification rules. Algorithms to construct and to match CR-tree are proposed. Inner-product is used to calculate the similarity between the category sub vector of the text and the category feature vector
and then is aggregated with the rules' confidence degree to serve as the foundation of text categorization. Experimental results show that the method presented achieves a micro-average value of categorization 92.42% with extracting only 30 feature terms
which is better than the results of AWOPR
KNN
and Bayes classifiers. And the time complexity of the method is the same as that of AWOPR
indicating that the cost to calculate both the similarity and the aggregation is acceptable.
LIU Bing, HSU W, MA Yiming. Integrating classification and association rule mining [C]∥Proceedings of the ACM International Conference on Knowledge Discovery and Data Mining. New York,USA: ACM, 1998: 80-86.
ZAÏANE O R, ANTONIE M L. Classifying text documents by associating terms with text categories [C]∥Proceedings of the 13th Australasian Database Conference. New York,USA: ACM, 2002: 215-222.
LI Wenmin, HAN Jiawei, PEI Jian. CMAR: Accurate and efficient classification based on multiple classification rules [C]∥Proceedings of the 2001 IEEE International Conference on Data Mining. Piscataway,NJ,USA: IEEE, 2001: 369-376.
陈晓云,陈袆,王雷,等.基于分类规则树的频繁模式文本分类[J].软件学报,2006,17(5):1017-1025.
CHEN Xiaoyun, CHEN Yi, WANG Lei, et al. Text categorization based on classification rules tree by frequent patterns [J]. Journal of Software, 2006, 17(5): 1017-1025.
陈晓云,胡运发. 基于自适应加权的文本关联分类[J].小型微型计算机系统,2007,28(1):116-121.
CHEN Xiaoyun, HU Yunfa. Text association categorization based on self-adaptive weighting [J]. Journal of Chinese Computer Systems, 2007, 28(1):116-121.
商炳章,白清源. 基于特征项权重改进的关联文本分类[J].计算机研究与发展,2008,45(S0):252-256.
SHANG Bingzhang, BAI Qingyuan. Improved association text classification based on feature weight [J]. Journal of Computer Research and Development, 2008, 45(S0): 252-256.
蔡金凤,白清源. 挖掘重要项集的关联文本分类[J].南京大学学报,2011,47(5): 544-550.
CAI Jinfeng, BAI Qingyuan. Association text classification of mining ItemSet significance [J]. Journal of Nanjing University, 2011, 47(5): 544-550.
BARALIS E, GARZA P. I-prune: item selection for associative classification [J]. International Journal of Intelligent Systems, 2012, 27(3): 279-299.
YANG Yiming, PEDERSON J O. A comparative study on feature selection in text categorization [C]∥Proceedings of the 14th International Conference on Machine Learning. San Francisco, CA, USA: Morgan Kaufmann, 1997: 412-420.
AGRAWAL R, SRIKANT R. Fast algorithms for mining association rules [C]∥Proceedings of the 20th VLDB Conference. San Francisco, CA, USA: Morgan Kaufmann, 1994: 487-499.
HAN Jiawei, PEI Jian, YIN Yiwen, et al. Mining frequent patterns without candidate generation: a frequent-pattern tree approach [J]. Data Mining and Knowledge Discovery, 2004, 8(1): 53-87.
SEBASTIANI F. Machine learning in automated text categorization [J]. ACM Computing Surveys, 2002, 34(1): 1-47.
杨攀,桂小林,田丰,等. 一种高效的用于话题检测的关键词元聚类方法. 2012,46(10):24-28.
霍战鹏,魏正英,张梦,等. 手机短信远程控制灌溉系统. 2012,46(10):36-41.
豆增发, 高琳. 利用膜粒子群优化和信息熵的医学文本特征选择. 2012,46(4):45-51.
薛峰,周亚东,高峰,等. 一种突发性热点话题在线发现与跟踪方法. 2011,45(12):64-69.
豆增发,高琳. 应用粒子群优化-条件随机域的文本生物实体识别. 2010,44(12):38-42.
赵煜,蔡皖东,樊娜,等. 采用并行遗传算法的文本分割研究. 2009,43(12):40-44.
姜庆民,吴宁,刘伟华. 面向入侵检测系统的模式匹配算法研究. 2009,43(2):58-62.
0
浏览量
4
下载量
4
CSCD
关联资源
相关文章
相关作者
相关机构
京公网安备11010802024621