西安交通大学计算机科学与技术系,西安,710049
网络首发:2009-08-10,
纸质出版:2009
移动端阅览
乔亚男, 齐勇, 侯迪. 具有立项过滤的信息检索查询词的分析方法[J]. 西安交通大学学报, 2009,43(8):6-10+63.
An Analysis Method of Query Terms for Information Retrieval Based on Isolated Terms Filtrating[J]. 2009, 43(8): 6-10+63.
针对传统查询词临近性(QTP)分析方法无法有效提高查准率的问题
提出了一种孤立项过滤的信息检索查询词分析方法.该方法根据词汇相似度较高的查询词对之间具有强可替代性这一事实
从查询词及其实例中分解出查询内的孤立项和文档内的孤立项
在分析查询词临近性之前预先进行孤立项过滤
使之不参与QTP统计量的计算
由此减小了过分强调临近性对查准率的影响.实验结果表明
对于词汇相似度差异比较显著的查询
进行孤立项过滤的查询词临近性分析方法的平均检索精确度比传统分析方法提高14%.
To address the issue that most of traditional query term proximity(QTP)statistics are lack of effective precision improvement
an analysis method of query terms based on isolated terms filtrating is proposed. The method is based on the fact that the query terms with higher similarity have higher substitutability. Isolated terms in-queries and isolated terms in-documents are extracted from query terms and their instances. The filtrating of isolated terms is earlier than the analyzing of QTP
so that the isolated terms are ignored in the calculating of QTP statistics and that the impact of overemphasis of QTP on precision is reduced. Experimental results on queries with significantly different proximities show that the performance of the QTP with isolated terms filtrating is higher(about 14%)than that of the QTP without isolated terms filtrating is.
TAO Tao, ZHAI Chengxiang. An exploration of proximity measures in information retrieval [C]∥Proceedings of the 30th Annual International ACM SIGIR Conference on Research and Development in Information Retrieval. New York, USA: ACM, 2007:295-302.
BUTTCHER S, CLARKE C, LUSHMAN B. Term proximity scoring for ad-hoc retrieval on very large text collections [C]∥Proceedings of the 26th Annual International ACM SIGIR Conference on Research and Development in Information Retrieval. New York, USA: ACM, 2006:621-622.
BUTTCHER S, CLARKE C L A. Efficiency vs. effectiveness in terabyte-scale information retrieval [C]∥Proceedings of the 14th Text Retrieval Conference(TREC). Gaithersburg, MD, USA: National Institute of Standards and Technology, 2005:1-6.
RASOLOFO Y, SAVOY J. Term proximity scoring for keyword-based retrieval systems [C]∥Proceedings of the 25th European Conference on IR Research. Heidelberg, Germany: Springer, 2003: 207-218.
PATWARDHAN S, PEDERSEN T. Using WordNet-based context vectors to estimate the semantic relatedness of concepts [C]∥Proceedings of EACL 2006 Workshop on Making Sense of Sense — Bringing Computational Linguistics and Psycholinguistics Together. Cambridge, MA, USA: MIT Press,2006:1-8.
JIANG J, CONRATH D. Semantic similarity based on corpus statistics and lexical taxonomy [C]∥Proceedings of International Conference of Research in Computational Linguistics. Heidelberg, Germany: Springer,1997: 1-15.
LIN D. An information-theoretic definition of similarity[C]∥Proceedings of International Conference on Machine Learning. San Francisco, USA: Morgan Kaufmann,1998: 296-304.
张云,冯博琴,麻首强,等. 蚁群-遗传融合的文本聚类算法[J]. 西安交通大学学报,2007,41(10): 1146-1150.
ZHANG Yun, FENG Boqin, MA Shouqiang, et al. Text clustering based on fusion of ant colony and genetic algorithms [J]. Journal of Xi'an Jiaotong University, 2007, 41(10): 1146-1150.
冯中慧,鲍军鹏,沈钧毅. 一种增量式文本软聚类算法[J]. 西安交通大学学报, 2007, 41(4): 398-401.
FENG Zhonghui, BAO Junpeng, SHEN Junyi. Incremental algorithm of text soft clustering [J]. Journal of Xi'an Jiaotong University, 2007, 41(4): 398-401.
LEWIS D D, YANG Y, ROSE T, et al. RCV1: a new benchmark collection for text categorization research [J]. Journal of Machine Learning Research, 2004, 5: 361-397.
ROBERTSON S E, WALKER S. Some simple effective approximations to the 2-poisson model for probabilistic weighted retrieval[C]∥Proceedings of the 17th Annual International ACM SIGIR Conference on Research and Development in Information Retrieval. Berlin, Germany: Springer-Verlag, 1994: 232-241.
0
浏览量
4
下载量
0
CSCD
关联资源
相关文章
相关作者
相关机构
京公网安备11010802024621