To address the issue that most of traditional query term proximity(QTP)statistics are lack of effective precision improvement
an analysis method of query terms based on isolated terms filtrating is proposed. The method is based on the fact that the query terms with higher similarity have higher substitutability. Isolated terms in-queries and isolated terms in-documents are extracted from query terms and their instances. The filtrating of isolated terms is earlier than the analyzing of QTP
so that the isolated terms are ignored in the calculating of QTP statistics and that the impact of overemphasis of QTP on precision is reduced. Experimental results on queries with significantly different proximities show that the performance of the QTP with isolated terms filtrating is higher(about 14%)than that of the QTP without isolated terms filtrating is.
关键词
Keywords
references
TAO Tao, ZHAI Chengxiang. An exploration of proximity measures in information retrieval [C]∥Proceedings of the 30th Annual International ACM SIGIR Conference on Research and Development in Information Retrieval. New York, USA: ACM, 2007:295-302.
BUTTCHER S, CLARKE C, LUSHMAN B. Term proximity scoring for ad-hoc retrieval on very large text collections [C]∥Proceedings of the 26th Annual International ACM SIGIR Conference on Research and Development in Information Retrieval. New York, USA: ACM, 2006:621-622.
BUTTCHER S, CLARKE C L A. Efficiency vs. effectiveness in terabyte-scale information retrieval [C]∥Proceedings of the 14th Text Retrieval Conference(TREC). Gaithersburg, MD, USA: National Institute of Standards and Technology, 2005:1-6.
RASOLOFO Y, SAVOY J. Term proximity scoring for keyword-based retrieval systems [C]∥Proceedings of the 25th European Conference on IR Research. Heidelberg, Germany: Springer, 2003: 207-218.
PATWARDHAN S, PEDERSEN T. Using WordNet-based context vectors to estimate the semantic relatedness of concepts [C]∥Proceedings of EACL 2006 Workshop on Making Sense of Sense — Bringing Computational Linguistics and Psycholinguistics Together. Cambridge, MA, USA: MIT Press,2006:1-8.
JIANG J, CONRATH D. Semantic similarity based on corpus statistics and lexical taxonomy [C]∥Proceedings of International Conference of Research in Computational Linguistics. Heidelberg, Germany: Springer,1997: 1-15.
LIN D. An information-theoretic definition of similarity[C]∥Proceedings of International Conference on Machine Learning. San Francisco, USA: Morgan Kaufmann,1998: 296-304.
ZHANG Yun, FENG Boqin, MA Shouqiang, et al. Text clustering based on fusion of ant colony and genetic algorithms [J]. Journal of Xi'an Jiaotong University, 2007, 41(10): 1146-1150.
FENG Zhonghui, BAO Junpeng, SHEN Junyi. Incremental algorithm of text soft clustering [J]. Journal of Xi'an Jiaotong University, 2007, 41(4): 398-401.
LEWIS D D, YANG Y, ROSE T, et al. RCV1: a new benchmark collection for text categorization research [J]. Journal of Machine Learning Research, 2004, 5: 361-397.
ROBERTSON S E, WALKER S. Some simple effective approximations to the 2-poisson model for probabilistic weighted retrieval[C]∥Proceedings of the 17th Annual International ACM SIGIR Conference on Research and Development in Information Retrieval. Berlin, Germany: Springer-Verlag, 1994: 232-241.