西安交通大学电子与信息工程学院,西安,710049
网络首发:2008-02-10,
纸质出版:2008
移动端阅览
崔舒宁, 朱丹军, 冯博琴, 等. 结合受控词汇表的生物基因本体标注与分类[J]. 西安交通大学学报, 2008,42(2):171-174.
崔舒宁, 朱丹军, 冯博琴, et al. Triage and Annotation of Biological Gene's Ontology Combined Controlled Glossary[J]. 2008, 42(2): 171-174.
通过研究有关基因的生物学文献特征
提出了一种能对生物基因文献进行自动标注与分类的方法.在K最邻近算法的基础上
采用了Chi-Square特征选择方案
并且在加权算法中突出了Chi-Square的选择特点.另外
采用文档逻辑分块法
将额外的生物受控词汇表中的信息所形成的向量直接引入到了分类算法中
以提高分类和标注的效果.实验表明
所提算法优于常用的单词频率/逆文档频率加权方法
其在文本检索大会(TREC)数据集上的分类、标注效果分别比TREC公布的最好结果提高了3.14%和4.12%.
Based on the K nearest neighbor algorithm
an improved method was proposed for selecting genes-related documents from biology literature
and then automatically annotating and classifying. The method employs the Chi-Square feature selection plan and highlights the Chi-Square selections in weighted calculations. Furthermore
the effect of classification and annotation was improved by dividing the documents into logical blocks and introducing additional vectors from biological resources MeSH into the classification algorithm directly. Experiment results show that the proposed method is better than the commonly used TFIDF(term frequency and inverse document frequency)weighting method
and the results tested on TREC(text retrieval conference)data sets are 3.14% higher in classification and 4.13% higher in annotation comparing to the best results announced TREC.
DIETRICH REBHOLZ-SCHUHMANN H K, COUTO F. Facts from text — is text mining ready to deliver ?[J]. PLoS Biology, 2005, 3(2):188-191.
ADITYA V P B, KINCAID R. An architecture for biological information extraction and representation [J]. Bioinformatics, 2005, 21(4):430-438.
NARAYANASWAMY M, RAVIKUMAR K E. A biological named entity recognizer [C]∥Proceedings of Pacific Symposium on Biocomputing. Hawaii, USA: World Scientific, 2003:427-438.
TSUJI J. Boosting precision and recall of dictionary-based protein name recognition [C]∥Proceedings of Atomic Level Characterizations. Hawaii, USA: Wiley Publisher, 2003:41-48.
ZHOU Guodong, SU Jian. Named entity recognition using an HMM-based chunk tagger [C]∥Proceedings of the 40th Annual Meeting of the Association for Computational Linguistics. San Francisco, USA: Morgan Kaufmann Publishers, 2002:473-480.
YANG Yiming, PEDERSEN J O. A comparative study on feature selection in text categorization [C]∥Proceedings of the 14th International Conference on Machine Learning. San Francisco, USA: Morgan Kaufmann Publishers, 1997:412-420.
ROSENFELD R.A maximum entropy approach to adaptive statistical language modeling [J].Computer, Speech, and Language,1996(10):187-228.
0
浏览量
5
下载量
2
CSCD
关联资源
相关文章
相关作者
相关机构
京公网安备11010802024621